An automated pipeline for construction of Reference Transcript Datasets (RTD) to enable rapid and accurate gene expression analysis in plant species
An automated pipeline for construction of Reference Transcript Datasets (RTD) to enable rapid and accurate gene expression analysis in plant species
批准号:
BB/S020160/1
负责人:
Runxuan Zhang
金额:
$40.38万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
基因是基因组上的基本物理和功能单位。基因在发育的不同时期关闭和开启,并对外部和内部信号做出反应。蛋白质编码基因被复制(转录)到前体信使RNA(前信使RNA)中,然后以不同的方式处理成mRNAs,然后再翻译成蛋白质。生物学研究的一个目标是通过测量基因表达的变化来了解基因是如何工作的。这是通过估计在任何特定时间或条件下产生的所有抄本的丰度来实现的。目前测量基因和转录物表达的技术称为RNA测序(RNA-SEQ),它通过对数百万转录本进行测序,使RNA水平能够在全基因组范围内被测量。两个主要的平台是Illumina和PacBio/Nanopore单分子测序,Illumina产生短片段(目前是75到250个碱基对),PacBio/Nanopore单分子测序产生全长转录片段。为了测量基因表达,Illumina短读数通常被映射到基因组并组装成转录本,这是一个不准确的过程。PacBio/Nanopore的测序错误率很高,不能产生足够深度的基因复盖。在化学和计算分析方面,这些技术继续快速发展,但这些平台的结合是目前产生RNA-SEQ数据的最佳方法。此外,最快和最准确的转录本和基因表达计算量化程序需要一个全面的转录本目录,我们称之为参考转录数据集(RTD)。在过去的四年里,我们开发了一个基于广泛的Illumina短阅读序列的拟南芥RTD(AtRTD2)。通过一系列迭代,我们开发了在删除虚假转录的同时识别和保留高置信度转录的计算方法。AtRTD2极大地提高了定量的准确性,例如,允许识别对寒冷反应的新的转录和剪接因子。现在的挑战是将这些知识和经验转化为其他植物和作物(以及动物)物种。目前,大多数植物物种的转录序列目录是不完整的,缺少大量的转录本,而对于那些拥有RNA-SEQ数据的物种,过时的分析程序产生了大量错误的转录本。通过开发AtRTD2,我们有了构建RTD的原型管道。关键功能是多个质量控制过滤器,可以删除错误组装的转录、多余的转录、嵌合体转录和转录片段。这些重复的多个步骤目前是单独编码的,虽然可以使用管道,但生成RTD将需要长达12个月的时间,并且需要生物信息学家的全职专业知识。我们将开发一种全自动管道(RTDBox),可供具有基本生物信息学技能的科学家或在转录组方面缺乏经验的生物信息学家使用。这种管道的设计还将允许通过自动纳入任何新的RNA-SEQ数据(Illumina、PacBio、Nanopore)来逐步改进RTD。在计划中,我们将开发一个转录评估套件(TES),它将提供评估指标,帮助生物学家识别和删除汇编程序中错误构建的转录,以及了解生成的RTD的质量和完整性。我们所有的经验和专业知识将汇集在一起,为植物科学家制作一个用户友好的软件,以便更准确地测量基因表达,从而改进全球生物过程的探索。
英文摘要
A gene is the basic physical and functional unit on the genome. Genes are turned off and on at different times of development and in response to external and internal signals. Protein-coding genes are copied (transcribed) into precursor messenger RNA (pre-mRNA) which are then processed in different ways into mRNAs which can then be translated into proteins. A goal of the biological research is to understand how genes work by measuring changes in gene expression. This is achieved by estimating the abundances of all of the transcripts produced at any particular time or condition. The current technologies to measure gene and transcript expression are called RNA sequencing (RNA-seq) which by sequencing millions of transcripts allows RNA levels to be measured on a genome-wide scale. The two main platforms are Illumina which generates short reads (currently 75 to 250 bp) and PacBio/Nanopore single molecule sequencing which produces full-length transcript reads. To measure gene expression, Illumina short reads are often mapped to the genome and assembled into transcripts which is an inaccurate process. PacBio/Nanopore have high sequencing error rates and do not generate sufficient depth of coverage of genes. These technologies, both in terms of chemistry and computational analyses, continue to advance at a rapid pace but a combination of the platforms is currently the best approach to generate RNA-seq data. In addition, the fastest and most accurate programs for computational quantification of transcript and gene expression require a comprehensive catalogue of transcripts which we call a Reference Transcript Dataset (RTD). Over the last four years, we developed an RTD for Arabidopsis (AtRTD2) based on extensive Illumina short read sequences. Through a series of iterations, we developed the computational methods to identify and retain high confidence transcripts while removing false transcripts. AtRTD2 greatly increased the accuracy of the quantification allowing, for example, identification of novel transcription and splicing factors in response to cold. The challenge now is to translate this knowledge and experience to other plant and crop (and animal) species. Currently, transcript sequence catalogues for most plant species are incomplete, missing large numbers of transcripts, and for those with RNA-seq data, out-of-date analysis procedures have produced large numbers of false transcripts. From developing AtRTD2, we have a prototype pipeline for constructing an RTD. The key features are multiple quality control filters which remove mis-assembled transcripts, redundant transcripts, chimaeric transcripts and transcript fragments. These multiple, iterative steps are currently individually coded and while the pipeline can be used, it will take up to 12 months to generate an RTD and requires the full-time expertise of a bioinformatician. We will develop a fully automated pipeline (RTDBox) which can be used by scientists with basic bioinformatics skills or bioinformaticians with little experience in transcriptomics. Such a pipeline would also be designed to allow the incremental improvement of the RTD with the automatic incorporation of any new RNA-seq data (Illumina, PacBio, Nanopore). Within the pipeline, we will develop a transcript evaluation suite (TES) which will provide evaluation metrics to help biologists to identify and remove mis-constructed transcripts from assembly programs as well as understand the quality and completeness of the RTD generated. All our experience and expertise will be brought together to make a user-friendly software for plant scientists to measure gene expressions more accurately and thereby improving the exploration of biological processes across the globe.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI:
10.26508/lsa.202101255
发表时间:
2022-08
期刊:
LIFE SCIENCE ALLIANCE
影响因子:
4.4
作者:
[Guo, Wenbin, Coulter, Max, Waugh, Robbie, Zhang, Runxuan]
通讯作者:
Zhang, Runxuan
DOI:
10.1186/s12864-019-6243-7
发表时间:
2019-12-11
期刊:
BMC GENOMICS
影响因子:
4.4
作者:
[Rapazote-Flores, Paulo, Bayer, Micha, Simpson, Craig G.]
通讯作者:
Simpson, Craig G.
DOI:
10.1080/15476286.2020.1858253
发表时间:
2021-11
期刊:
RNA biology
影响因子:
4.1
作者:
[Guo W, Tzioutziou NA, Stephen G, Milne I, Calixto CP, Waugh R, Brown JWS, Zhang R]
通讯作者:
Zhang R
国内基金
海外基金
FAST连续观测数据处理的pipeline开发
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
云台40米射电望远镜脉冲星后端实时pipeline关键技术研究
-
批准号:12063003
-
项目类别:地区科学基金项目
-
资助金额:37.0万元
-
批准年份:2020
-
负责人:戴伟
-
依托单位:
FAST高性能Pipeline关键技术研究
-
批准号:U1731125
-
项目类别:联合基金项目
-
资助金额:46.0万元
-
批准年份:2017
-
负责人:肖健
-
依托单位: