An automated pipeline for construction of Reference Transcript Datasets (RTD) to enable rapid and accurate gene expression analysis in plant species
An automated pipeline for construction of Reference Transcript Datasets (RTD) to enable rapid and accurate gene expression analysis in plant species
批准号:
BB/S020160/1
负责人:
Runxuan Zhang
金额:
$40.38万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
A gene is the basic physical and functional unit on the genome. Genes are turned off and on at different times of development and in response to external and internal signals. Protein-coding genes are copied (transcribed) into precursor messenger RNA (pre-mRNA) which are then processed in different ways into mRNAs which can then be translated into proteins. A goal of the biological research is to understand how genes work by measuring changes in gene expression. This is achieved by estimating the abundances of all of the transcripts produced at any particular time or condition. The current technologies to measure gene and transcript expression are called RNA sequencing (RNA-seq) which by sequencing millions of transcripts allows RNA levels to be measured on a genome-wide scale. The two main platforms are Illumina which generates short reads (currently 75 to 250 bp) and PacBio/Nanopore single molecule sequencing which produces full-length transcript reads. To measure gene expression, Illumina short reads are often mapped to the genome and assembled into transcripts which is an inaccurate process. PacBio/Nanopore have high sequencing error rates and do not generate sufficient depth of coverage of genes. These technologies, both in terms of chemistry and computational analyses, continue to advance at a rapid pace but a combination of the platforms is currently the best approach to generate RNA-seq data. In addition, the fastest and most accurate programs for computational quantification of transcript and gene expression require a comprehensive catalogue of transcripts which we call a Reference Transcript Dataset (RTD). Over the last four years, we developed an RTD for Arabidopsis (AtRTD2) based on extensive Illumina short read sequences. Through a series of iterations, we developed the computational methods to identify and retain high confidence transcripts while removing false transcripts. AtRTD2 greatly increased the accuracy of the quantification allowing, for example, identification of novel transcription and splicing factors in response to cold. The challenge now is to translate this knowledge and experience to other plant and crop (and animal) species. Currently, transcript sequence catalogues for most plant species are incomplete, missing large numbers of transcripts, and for those with RNA-seq data, out-of-date analysis procedures have produced large numbers of false transcripts. From developing AtRTD2, we have a prototype pipeline for constructing an RTD. The key features are multiple quality control filters which remove mis-assembled transcripts, redundant transcripts, chimaeric transcripts and transcript fragments. These multiple, iterative steps are currently individually coded and while the pipeline can be used, it will take up to 12 months to generate an RTD and requires the full-time expertise of a bioinformatician. We will develop a fully automated pipeline (RTDBox) which can be used by scientists with basic bioinformatics skills or bioinformaticians with little experience in transcriptomics. Such a pipeline would also be designed to allow the incremental improvement of the RTD with the automatic incorporation of any new RNA-seq data (Illumina, PacBio, Nanopore). Within the pipeline, we will develop a transcript evaluation suite (TES) which will provide evaluation metrics to help biologists to identify and remove mis-constructed transcripts from assembly programs as well as understand the quality and completeness of the RTD generated. All our experience and expertise will be brought together to make a user-friendly software for plant scientists to measure gene expressions more accurately and thereby improving the exploration of biological processes across the globe.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI:
10.26508/lsa.202101255
发表时间:
2022-08
期刊:
LIFE SCIENCE ALLIANCE
影响因子:
4.4
作者:
[Guo, Wenbin, Coulter, Max, Waugh, Robbie, Zhang, Runxuan]
通讯作者:
Zhang, Runxuan
DOI:
10.1186/s12864-019-6243-7
发表时间:
2019-12-11
期刊:
BMC GENOMICS
影响因子:
4.4
作者:
[Rapazote-Flores, Paulo, Bayer, Micha, Simpson, Craig G.]
通讯作者:
Simpson, Craig G.
DOI:
10.1080/15476286.2020.1858253
发表时间:
2021-11
期刊:
RNA biology
影响因子:
4.1
作者:
[Guo W, Tzioutziou NA, Stephen G, Milne I, Calixto CP, Waugh R, Brown JWS, Zhang R]
通讯作者:
Zhang R
国内基金
海外基金
FAST连续观测数据处理的pipeline开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
云台40米射电望远镜脉冲星后端实时pipeline关键技术研究
-
批准号:12063003
-
项目类别:地区科学基金项目
-
资助金额:37.0万元
-
批准年份:2020
-
负责人:戴伟
-
依托单位:
FAST高性能Pipeline关键技术研究
-
批准号:U1731125
-
项目类别:联合基金项目
-
资助金额:46.0万元
-
批准年份:2017
-
负责人:肖健
-
依托单位: