Benchmarking Next-Generation Transcriptome Sequencing for Functional and Evolutionary Genomics

Benchmarking Next-Generation Transcriptome Sequencing for Functional and Evolutionary Genomics
复制标题

DOI:
10.1093/molbev/msp188
复制
发表时间:
2009-12-01
影响因子:
10.7
通讯作者:
Rokas, Antonis
Rokas, Antonis
中科院分区:
生物学1区
文献类型:
--
作者:
Gibbons, John G.;Janson, Eric M.;Rokas, Antonis

文献摘要

被引文献

相似文献

下一代测序为非模式生物的基因组分析打开了大门。产生长序列读出(200-400bp)的技术越来越多地用于非模式生物的进化研究,但可以以较低成本产生的短序列读出(30-50bp)被认为对从头测序应用的效用有限。在这里,我们通过对热带病媒介埃及伊蚊和冈比亚按蚊的转录本进行短读测序来检验这一假设,这两种媒介的完整基因组序列都是可用的。将我们的结果与参考基因组进行比较,使我们能够准确地评估我们的“测试”数据的数量、质量以及功能和进化信息内容。我们为每个物种产生了超过7亿个核苷酸测序数据,这些序列数据组装成超过21,000个测试重叠群,每个物种超过100个核苷酸,覆盖了伊蚊参考转录组的27%。值得注意的是,测试重叠群中的替换错误率类似于每个站点0.25%,很少有indels或组装错误。这两个物种的测试重叠群富含涉及能量生产和蛋白质合成的基因,而在涉及转录和分化的基因中表达不足。使用测试重叠群进行的直射同源预测在数亿年的进化过程中是准确的。我们的结果证明了短读转录组测序在非模式生物基因组研究中的相当大的实用性,并为进化研究提供了一种评估下一代数据信息含量的方法。
Next-generation sequencing has opened the door to genomic analysis of nonmodel organisms. Technologies generating long-sequence reads (200-400 bp) are increasingly used in evolutionary studies of nonmodel organisms, but the short-sequence reads (30-50 bp) that can be produced at lower cost are thought to be of limited utility for de novo sequencing applications. Here, we tested this assumption by short-read sequencing the transcriptomes of the tropical disease vectors Aedes aegypti and Anopheles gambiae, for which complete genome sequences are available. Comparison of our results to the reference genomes allowed us to accurately evaluate the quantity, quality, and functional and evolutionary information content of our "test" data. We produced more than 0.7 billion nucleotides of sequenced data per species that assembled into more than 21,000 test contigs larger than 100 bp per species and covered similar to 27% of the Aedes reference transcriptome. Remarkably, the substitution error rate in the test contigs was similar to 0.25% per site, with very few indels or assembly errors. Test contigs of both species were enriched for genes involved in energy production and protein synthesis and underrepresented in genes involved in transcription and differentiation. Ortholog prediction using the test contigs was accurate across hundreds of millions of years of evolution. Our results demonstrate the considerable utility of short-read transcriptome sequencing for genomic studies of nonmodel organisms and suggest an approach for assessing the information content of next-generation data for evolutionary studies.