Comparative performance of transcriptome assembly methods for non-model organisms.

Comparative performance of transcriptome assembly methods for non-model organisms.
复制标题

DOI:
10.1186/s12864-016-2923-8
复制
发表时间:
2016-07-27
期刊:
影响因子:
4.4
通讯作者:
Armbruster PA
Armbruster PA
中科院分区:
生物学2区
文献类型:
--
作者:
Huang X;Chen XG;Armbruster PA

文献摘要

被引文献

相似文献

下一代测序的技术革命为在基因组或转录组水平上研究任何感兴趣的生物体带来了前所未有的机会。转录组组装是使用RNA测序(RNA-Seq)研究感兴趣表型的分子基础的关键第一步。然而,组装大量短RNA-Seq读数的最佳策略仍然没有得到解决,特别是对于没有测序基因组的生物体。本研究比较了四种转录组组装方法,包括广泛使用的从头组装器(Trinity),两种转录组重组策略,利用来自密切相关物种的蛋白质组和基因组资源(基于参考的重组和TransPS)和基因组指导的组装器(Cufflinks)。这四个装配策略进行了比较,使用一个全面的转录组数据库白纹伊蚊,基因组序列最近已经完成。通过生成的重叠群的数量、重叠群长度分布、配对末端读段作图百分比和通过BLASTX的基因模型表示来评估各种组装体的质量。我们的研究结果表明,从头组装产生类似数量的基因模型相对于基因组引导的组装与一个片段化的参考,但产生最高水平的冗余,并需要最多的计算能力。使用密切相关的参考基因组来指导转录组组装可以产生有偏差的重叠群序列。增加转录组组装中使用的读段的数量倾向于增加组装内的冗余,并降低重叠群和参考蛋白质序列之间的中值重叠群长度和同一性百分比。这项研究为来自具有或不具有测序基因组的生物体的RNA-Seq数据的转录组组装提供了一般指导。最佳的转录组组装策略将取决于后续的下游分析。然而,我们的结果强调了从头组装的功效,当参考基因组组装被片段化时,从头组装可以与基因组引导组装一样有效。如果基因组组装和足够的计算资源可用,则将联合收割机从头组装和基因组引导的组装组合可能是有益的。当使用密切相关的参考基因组来指导转录组组装时应谨慎。转录组组装中使用的读取对的数量不一定与组装的质量相关。本文的在线版本(doi:10.1186/s12864-016-2923-8)包含补充材料,可供授权用户使用。
The technological revolution in next-generation sequencing has brought unprecedented opportunities to study any organism of interest at the genomic or transcriptomic level. Transcriptome assembly is a crucial first step for studying the molecular basis of phenotypes of interest using RNA-Sequencing (RNA-Seq). However, the optimal strategy for assembling vast amounts of short RNA-Seq reads remains unresolved, especially for organisms without a sequenced genome. This study compared four transcriptome assembly methods, including a widely used de novo assembler (Trinity), two transcriptome re-assembly strategies utilizing proteomic and genomic resources from closely related species (reference-based re-assembly and TransPS) and a genome-guided assembler (Cufflinks). These four assembly strategies were compared using a comprehensive transcriptomic database of Aedes albopictus, for which a genome sequence has recently been completed. The quality of the various assemblies was assessed by the number of contigs generated, contig length distribution, percent paired-end read mapping, and gene model representation via BLASTX. Our results reveal that de novo assembly generates a similar number of gene models relative to genome-guided assembly with a fragmented reference, but produces the highest level of redundancy and requires the most computational power. Using a closely related reference genome to guide transcriptome assembly can generate biased contig sequences. Increasing the number of reads used in the transcriptome assembly tends to increase the redundancy within the assembly and decrease both median contig length and percent identity between contigs and reference protein sequences. This study provides general guidance for transcriptome assembly of RNA-Seq data from organisms with or without a sequenced genome. The optimal transcriptome assembly strategy will depend upon the subsequent downstream analyses. However, our results emphasize the efficacy of de novo assembly, which can be as effective as genome-guided assembly when the reference genome assembly is fragmented. If a genome assembly and sufficient computational resources are available, it can be beneficial to combine de novo and genome-guided assemblies. Caution should be taken when using a closely related reference genome to guide transcriptome assembly. The quantity of read pairs used in the transcriptome assembly does not necessarily correlate with the quality of the assembly. The online version of this article (doi:10.1186/s12864-016-2923-8) contains supplementary material, which is available to authorized users.