De novo transcriptome assembly: A comprehensive cross-species comparison of short-read RNA-Seq assemblers

De novo transcriptome assembly: A comprehensive cross-species comparison of short-read RNA-Seq assemblers
复制标题

DOI:
10.1093/gigascience/giz039
复制
发表时间:
2019-05-01
期刊:
影响因子:
9.2
通讯作者:
Marz, Manja
Marz, Manja
中科院分区:
生物学2区
文献类型:
--
作者:
Hoelzer, Martin;Marz, Manja

文献摘要

被引文献

相似文献

背景资料:近年来,大规模并行互补DNA测序(RNA测序[RNA-Seq])已成为一种快速,成本效益高,强大的技术,以各种方式研究整个转录组。特别是,对于非模式生物和在缺乏适当的参考基因组的情况下,RNA-Seq用于从头重建转录组。虽然最近非模式生物的从头转录组组装一直在上升,新的工具也在不断开发,但关于应该使用哪种组装软件来构建全面的从头组装,仍然存在知识差距。结果如下:在这里,我们提出了一个大规模的比较研究,其中10从头组装工具应用于9个RNA-Seq数据集跨越不同的生命王国。总的来说,我们构建了超过200个单一组件,并在20个基于生物学和无参考指标的组合上评估了它们的性能。我们的研究是伴随着一个全面的和可扩展的电子补充,总结了所有的数据集,装配执行指令和评估结果。Trinity、SPAdes和Trans-ABySS,其次是布里杰和SOAPdenovo-Trans,总体上优于所比较的其他工具。此外,我们观察到每个汇编程序的性能存在物种特异性差异。没有一种工具能够为所有数据集提供最佳结果。结论:我们建议仔细选择和标准化的评价指标,以选择最佳的组装结果作为一个全面的从头转录组组装重建的关键步骤。
Background: In recent years, massively parallel complementary DNA sequencing (RNA sequencing [RNA-Seq]) has emerged as a fast, cost-effective, and robust technology to study entire transcriptomes in various manners. In particular, for non-model organisms and in the absence of an appropriate reference genome, RNA-Seq is used to reconstruct the transcriptome de novo. Although the de novo transcriptome assembly of non-model organisms has been on the rise recently and new tools are frequently developing, there is still a knowledge gap about which assembly software should be used to build a comprehensive de novo assembly. Results: Here, we present a large-scale comparative study in which 10 de novo assembly tools are applied to 9 RNA-Seq data sets spanning different kingdoms of life. Overall, we built >200 single assemblies and evaluated their performance on a combination of 20 biological-based and reference-free metrics. Our study is accompanied by a comprehensive and extensible Electronic Supplement that summarizes all data sets, assembly execution instructions, and evaluation results. Trinity, SPAdes, and Trans-ABySS, followed by Bridger and SOAPdenovo-Trans, generally outperformed the other tools compared. Moreover, we observed species-specific differences in the performance of each assembler. No tool delivered the best results for all data sets. Conclusions: We recommend a careful choice and normalization of evaluation metrics to select the best assembling results as a critical step in the reconstruction of a comprehensive de novo transcriptome assembly.