Single-molecule real-time transcript sequencing facilitates common wheat genome annotation and grain transcriptome research.

Single-molecule real-time transcript sequencing facilitates common wheat genome annotation and grain transcriptome research.
复制标题

单分子实时转录测序促进常见小麦基因组注释和谷物转录组研究

DOI:
10.1186/s12864-015-2257-y
复制
发表时间:
2015-12-09
期刊:
影响因子:
4.4
通讯作者:
Wang D
Wang D
中科院分区:
生物学2区
文献类型:
--
作者:
Dong L;Liu H;Zhang J;Yang S;Kong G;Chu JS;Chen N;Wang D

文献摘要

被引文献

相似文献

普通小麦(Triticum aestivum,AABBDD)庞大而复杂的六倍体基因组极大地阻碍了其基因组学研究。在这里,我们研究了转录在普通小麦发育颖果使用新兴的单分子实时(SMRT)测序技术PacBio RSII,并评估所得的数据,以提高普通小麦基因组注释和谷物转录组研究。我们获得了197,709个全长非嵌合(FLNC)读段,估计其中74.6%携带完整的开放阅读框架。总共鉴定了91,881个高质量FLNC读数,并将其定位到16,188个染色体位点,对应于13,162个已知基因和3026个先前未注释的新基因。尽管一些FLNC读段不能明确地映射到当前的基因组序列草图,但它们中的许多可能用于研究高度相似的同源或旁系同源基因座或用于在进一步研究中改进染色体重叠群组装。91,881个高质量FLNC读数代表了22,768个独特的转录本,其中9591个是新发现的。我们发现了180个转录本,每个转录本跨越两个或三个先前注释的相邻位点,这表明它们应该合并以形成正确的基因模型。最后,我们的数据有助于识别6030个基因的差异调节颖果发育过程中,和全长转录的72个面筋基因的成员,是重要的普通小麦的最终使用质量控制。我们的工作证明了PacBio转录测序的价值,提高普通小麦基因组注释,通过发现基因座和全长转录本没有发现以前。该资源为小麦结构基因组学和籽粒转录组学的进一步研究提供了参考。本文的在线版本(doi:10.1186/s12864-015-2257-y)包含补充材料,可供授权用户使用。
The large and complex hexaploid genome has greatly hindered genomics studies of common wheat (Triticum aestivum, AABBDD). Here, we investigated transcripts in common wheat developing caryopses using the emerging single-molecule real-time (SMRT) sequencing technology PacBio RSII, and assessed the resultant data for improving common wheat genome annotation and grain transcriptome research. We obtained 197,709 full-length non-chimeric (FLNC) reads, 74.6 % of which were estimated to carry complete open reading frame. A total of 91,881 high-quality FLNC reads were identified and mapped to 16,188 chromosomal loci, corresponding to 13,162 known genes and 3026 new genes not annotated previously. Although some FLNC reads could not be unambiguously mapped to the current draft genome sequence, many of them are likely useful for studying highly similar homoeologous or paralogous loci or for improving chromosomal contig assembly in further research. The 91,881 high-quality FLNC reads represented 22,768 unique transcripts, 9591 of which were newly discovered. We found 180 transcripts each spanning two or three previously annotated adjacent loci, suggesting that they should be merged to form correct gene models. Finally, our data facilitated the identification of 6030 genes differentially regulated during caryopsis development, and full-length transcripts for 72 transcribed gluten gene members that are important for the end-use quality control of common wheat. Our work demonstrated the value of PacBio transcript sequencing for improving common wheat genome annotation through uncovering the loci and full-length transcripts not discovered previously. The resource obtained may aid further structural genomics and grain transcriptome studies of common wheat. The online version of this article (doi:10.1186/s12864-015-2257-y) contains supplementary material, which is available to authorized users.