A survey of the sorghum transcriptome using single-molecule long reads.

A survey of the sorghum transcriptome using single-molecule long reads.
复制标题

DOI:
10.1038/ncomms11706
复制
发表时间:
2016-06-24
影响因子:
16.6
通讯作者:
Reddy AS
Reddy AS
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Abdel-Ghany SE;Hamilton M;Jacobi JL;Ngam P;Devitt N;Schilkey F;Ben-Hur A;Reddy AS

文献摘要

被引文献

相似文献

前体mRNA的选择性剪接和选择性聚腺苷酸化(阿帕)对真核生物转录组多样性、基因组编码能力和基因调控机制有重要作用。第二代测序技术已被广泛用于分析转录组。然而,短读数据的主要限制是难以准确预测全长剪接异构体。在这里,我们使用Pacific Biosciences单分子实时长读同种型测序对高粱转录组进行测序,并开发了一个名为TAPIS的管道(用于同种型测序的转录组分析管道),以识别全长剪接同种型和阿帕位点。我们的分析揭示了转录组范围内的全长亚型在前所未有的规模超过11,000新的剪接亚型。此外,我们还发现了约11,000个表达基因和2,100多个新基因的阿帕。这些结果极大地增强了高粱基因注释,并有助于研究这种重要生物能源作物的基因调控。TAPIS管道将作为分析任何生物体的Iso-Seq数据的有用工具。选择性剪接和选择性聚腺苷酸化(阿帕)有助于mRNA多样性,但难以使用短读RNA-seq数据进行评估。在这里,作者使用单分子长读同种型测序,并开发了一个计算管道来识别高粱中的全长剪接同种型和阿帕位点。
Alternative splicing and alternative polyadenylation (APA) of pre-mRNAs greatly contribute to transcriptome diversity, coding capacity of a genome and gene regulatory mechanisms in eukaryotes. Second-generation sequencing technologies have been extensively used to analyse transcriptomes. However, a major limitation of short-read data is that it is difficult to accurately predict full-length splice isoforms. Here we sequenced the sorghum transcriptome using Pacific Biosciences single-molecule real-time long-read isoform sequencing and developed a pipeline called TAPIS (Transcriptome Analysis Pipeline for Isoform Sequencing) to identify full-length splice isoforms and APA sites. Our analysis reveals transcriptome-wide full-length isoforms at an unprecedented scale with over 11,000 novel splice isoforms. Additionally, we uncover APA of ∼11,000 expressed genes and more than 2,100 novel genes. These results greatly enhance sorghum gene annotations and aid in studying gene regulation in this important bioenergy crop. The TAPIS pipeline will serve as a useful tool to analyse Iso-Seq data from any organism. Alternative splicing and alternative polyadenylation (APA) contribute to mRNA diversity but are difficult to assess using short read RNA-seq data. Here, the authors use single molecule long-read isoform sequencing and develop a computational pipeline to identify full-length splice isoforms and APA sites in sorghum.