A survey of the complex transcriptome from the highly polyploid sugarcane genome using full-length isoform sequencing and de novo assembly from short read sequencing.

A survey of the complex transcriptome from the highly polyploid sugarcane genome using full-length isoform sequencing and de novo assembly from short read sequencing.
复制标题

DOI:
10.1186/s12864-017-3757-8
复制
发表时间:
2017-05-22
期刊:
影响因子:
4.4
通讯作者:
Henry RJ
Henry RJ
中科院分区:
生物学2区
文献类型:
--
作者:
Hoang NV;Furtado A;Mason PJ;Marquardt A;Kasirajan L;Thirugnanasambandam PP;Botha FC;Henry RJ

文献摘要

被引文献

相似文献

尽管甘蔗在糖和生物能源生产中具有重要的经济意义,但目前还没有一个参考基因组。大多数甘蔗转录组学研究都是基于甘蔗基因索引(SoGI),表达序列标签(EST)和从头组装的短读段转录重叠群;因此,甘蔗转录组的知识是有限的,与转录本的长度和数量的转录异构体。使用PacBio同种型测序(Iso-Seq)对来自22个品种的不同发育阶段的叶、节间和根组织的合并RNA样品进行甘蔗转录组测序,以探索捕获全长转录物同种型的潜力。共获得107,598个独特的转录异构体,占预测的甘蔗基因总数的约71%。该数据集的大部分(92%)与植物蛋白质数据库相匹配,而超过2%是新的转录本,超过2%是推定的长非编码RNA。约56%和23%的总序列分别对基因本体和KEGG途径数据库进行了注释。与来自相同实验和公共数据库的节间样品的Illumina RNA测序(RNA-Seq)的从头重叠群的比较显示,Iso-Seq方法回收了更多全长转录物同种型,具有更高的N50和最大1,000个蛋白质的平均长度;而在RNA-Seq中捕获了基因含量和RNA多样性的更大代表性。只有62%的PacBio转录异构体与67%的从头重叠群匹配,而不匹配的比例分别归因于PacBio中包含叶/根组织和标准化,以及从头组装中更多基因内容和RNA类别的代表。约69%的PacBio转录异构体和41%的从头重叠群与高粱基因组比对,表明两个基因组的基因区域中的直系同源物高度保守。转录组数据集应有助于改进甘蔗基因模型和甘蔗蛋白质预测,并将作为甘蔗转录表达分析的参考数据库。本文的在线版本(doi:10.1186/s12864-017-3757-8)包含补充材料,可供授权用户使用。
Despite the economic importance of sugarcane in sugar and bioenergy production, there is not yet a reference genome available. Most of the sugarcane transcriptomic studies have been based on Saccharum officinarum gene indices (SoGI), expressed sequence tags (ESTs) and de novo assembled transcript contigs from short-reads; hence knowledge of the sugarcane transcriptome is limited in relation to transcript length and number of transcript isoforms. The sugarcane transcriptome was sequenced using PacBio isoform sequencing (Iso-Seq) of a pooled RNA sample derived from leaf, internode and root tissues, of different developmental stages, from 22 varieties, to explore the potential for capturing full-length transcript isoforms. A total of 107,598 unique transcript isoforms were obtained, representing about 71% of the total number of predicted sugarcane genes. The majority of this dataset (92%) matched the plant protein database, while just over 2% was novel transcripts, and over 2% was putative long non-coding RNAs. About 56% and 23% of total sequences were annotated against the gene ontology and KEGG pathway databases, respectively. Comparison with de novo contigs from Illumina RNA-Sequencing (RNA-Seq) of the internode samples from the same experiment and public databases showed that the Iso-Seq method recovered more full-length transcript isoforms, had a higher N50 and average length of largest 1,000 proteins; whereas a greater representation of the gene content and RNA diversity was captured in RNA-Seq. Only 62% of PacBio transcript isoforms matched 67% of de novo contigs, while the non-matched proportions were attributed to the inclusion of leaf/root tissues and the normalization in PacBio, and the representation of more gene content and RNA classes in the de novo assembly, respectively. About 69% of PacBio transcript isoforms and 41% of de novo contigs aligned with the sorghum genome, indicating the high conservation of orthologs in the genic regions of the two genomes. The transcriptome dataset should contribute to improved sugarcane gene models and sugarcane protein predictions; and will serve as a reference database for analysis of transcript expression in sugarcane. The online version of this article (doi:10.1186/s12864-017-3757-8) contains supplementary material, which is available to authorized users.