A comparison across non-model animals suggests an optimal sequencing depth for de novo transcriptome assembly.

A comparison across non-model animals suggests an optimal sequencing depth for de novo transcriptome assembly.
复制标题

DOI:
10.1186/1471-2164-14-167
复制
发表时间:
2013-03-12
期刊:
影响因子:
4.4
通讯作者:
Haddock SH
Haddock SH
中科院分区:
生物学2区
文献类型:
--
作者:
Francis WR;Christianson LM;Kiko R;Powers ML;Shaner NC;Haddock SH

文献摘要

参考文献

被引文献

相似文献

基因组资源的缺乏可能对非模式生物的研究提出挑战。转录组测序提供了一种有吸引力的方法来收集有关基因和基因表达的信息,而不需要参考基因组。然而,目前还不清楚什么样的测序深度足以组装转录组从头用于这些目的。我们组装转录组的动物从六个不同的门(环节动物,节肢动物,脊索动物,刺胞动物,Ctenophores和软体动物)在定期增量的读取天鹅绒/绿洲和Trinity,以确定如何读取计数影响大会。这包括小鼠心脏读数的组装,因为我们可以将其与可用的参考基因组进行比较。我们发现,在整个动物与组织的组件的质量差异。随着读段的增加,整个动物组装显示转录本的快速增加和保守基因的发现,而单组织组装显示保守基因的发现较慢,尽管组装的转录本通常更长。对小鼠组装体的更深入检查显示,随着更多的读取,组装错误变得更频繁,但可以用更严格的组装参数来减轻这种错误。这些组装趋势表明,对于RNA水平覆盖,产生的代表性组装体具有少至2000万个组织样品读数和3000万个全动物读数。这些深度在覆盖范围和噪声之间提供了良好的平衡。超过6000万次读取,新基因的发现率很低,高表达基因的测序错误可能会积累。最后,管水母(多态刺胞动物)是一个例外,可能需要替代组装策略。
The lack of genomic resources can present challenges for studies of non-model organisms. Transcriptome sequencing offers an attractive method to gather information about genes and gene expression without the need for a reference genome. However, it is unclear what sequencing depth is adequate to assemble the transcriptome de novo for these purposes. We assembled transcriptomes of animals from six different phyla (Annelids, Arthropods, Chordates, Cnidarians, Ctenophores, and Molluscs) at regular increments of reads using Velvet/Oases and Trinity to determine how read count affects the assembly. This included an assembly of mouse heart reads because we could compare those against the reference genome that is available. We found qualitative differences in the assemblies of whole-animals versus tissues. With increasing reads, whole-animal assemblies show rapid increase of transcripts and discovery of conserved genes, while single-tissue assemblies show a slower discovery of conserved genes though the assembled transcripts were often longer. A deeper examination of the mouse assemblies shows that with more reads, assembly errors become more frequent but such errors can be mitigated with more stringent assembly parameters. These assembly trends suggest that representative assemblies are generated with as few as 20 million reads for tissue samples and 30 million reads for whole-animals for RNA-level coverage. These depths provide a good balance between coverage and noise. Beyond 60 million reads, the discovery of new genes is low and sequencing errors of highly-expressed genes are likely to accumulate. Finally, siphonophores (polymorphic Cnidarians) are an exception and possibly require alternate assembly strategies.
DOI: 10.1093/nar/gkn916
发表时间: 2009-01
影响因子: 14.9
作者:
Parra G;Bradnam K;Ning Z;Keane T;Korf I
通讯作者: Korf I
DOI: 10.1186/1471-2164-12-317
发表时间: 2011-06-16
期刊: BMC genomics
影响因子: 4.4
作者:
Feldmeyer B;Wheat CW;Krezdorn N;Rotter B;Pfenninger M
通讯作者: Pfenninger M
DOI: 10.1038/nmeth.1223
发表时间: 2008-07-01
期刊: NATURE METHODS
影响因子: 48
作者:
Cloonan, Nicole;Forrest, Alistair R. R.;Grimmond, Sean M.
通讯作者: Grimmond, Sean M.
DOI: 10.1186/1471-2164-10-203
发表时间: 2009-04-29
期刊: BMC genomics
影响因子: 4.4
作者:
Hale MC;McCormick CR;Jackson JR;Dewoody JA
通讯作者: Dewoody JA
DOI: 10.1101/gr.104372.109
发表时间: 2010-08-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Romiguier, Jonathan;Ranwez, Vincent;Galtier, Nicolas
通讯作者: Galtier, Nicolas