Comparative analysis of de novo transcriptome assembly.

Comparative analysis of de novo transcriptome assembly.
复制标题

从头转录组组件的比较分析。

DOI:
10.1007/s11427-013-4444-x
复制
发表时间:
2013-03
期刊:
Science China. Life sciences
影响因子:
--
通讯作者:
Zhang KK
Zhang KK
中科院分区:
其他
文献类型:
--
作者:
Clarke K;Yang Y;Marsh R;Xie L;Zhang KK

文献摘要

参考文献

被引文献

相似文献

新一代测序技术的快速发展对数据处理和分析提出了重大的计算挑战。de Bruijn图是一种快速的DNA重组算法,它已被成功地应用于基因组DNA的从头组装,但其在转录组组装中的性能还不清楚。在这项研究中,我们使用模拟和真实的RNA-Seq数据,从人工RNA模板或人类转录本,以评估五个从头组装,ABySS,米拉,Trinity,天鹅绒和绿洲。在这些汇编器中,ABySS、Trinity、Velvet和Oases都是基于de Bruijn图的,Mira使用重叠图算法。从外部RNA对照联盟(ERCC)数据和人22号染色体中选择各种数量的RNA短读段。然后计算每个组装程序产生的重叠群的许多统计数据。每个实验重复多次以获得平均统计量和标准误差估计值。Trinity对ERCC和人类数据都具有相对良好的性能,但它可能无法始终生成全长转录本。ABySS是最快的方法,但其组装质量低。Mira在人类22号染色体上的重叠群定位率较高,但计算速度不理想。我们的研究结果表明,转录本组装仍然是生物信息学社会的挑战性问题。因此,需要一种新的组装器来组装由下一代测序技术产生的转录组数据。
The fast development of next-generation sequencing technology presents a major computational challenge for data processing and analysis. A fast algorithm, de Bruijn graph has been successfully used for genome DNA de novo assembly; nevertheless, its performance for transcriptome assembly is unclear. In this study, we used both simulated and real RNA-Seq data, from either artificial RNA templates or human transcripts, to evaluate five de novo assemblers, ABySS, Mira, Trinity, Velvet and Oases. Of these assemblers, ABySS, Trinity, Velvet and Oases are all based on de Bruijn graph, and Mira uses an overlap graph algorithm. Various numbers of RNA short reads were selected from the External RNA Control Consortium (ERCC) data and human chromosome 22. A number of statistics were then calculated for the resulting contigs from each assembler. Each experiment was repeated multiple times to obtain the mean statistics and standard error estimate. Trinity had relative good performance for both ERCC and human data, but it may not consistently generate full length transcripts. ABySS was the fastest method but its assembly quality was low. Mira gave a good rate for mapping its contigs onto human chromosome 22, but its computational speed is not satisfactory. Our results suggest that transcript assembly remains a challenge problem for bioinformatics society. Therefore, a novel assembler is in need for assembling transcriptome data generated by next generation sequencing technique.
DOI: 10.1093/bioinformatics/bts094
发表时间: 2012-04-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Schulz MH;Zerbino DR;Vingron M;Birney E
通讯作者: Birney E
DOI: 10.1089/cmb.1995.2.291
发表时间: 1995-01-01
期刊: Journal of computational biology : a journal of computational molecular cell biology
影响因子: --
作者:
Idury, R M;Waterman, M S
通讯作者: Waterman, M S
DOI: 10.1101/gr.731003
发表时间: 2003-01-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Mullikin, JC;Ning, ZM
通讯作者: Ning, ZM
DOI: 10.1186/gb-2009-10-3-r25
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Langmead B;Trapnell C;Pop M;Salzberg SL
通讯作者: Salzberg SL
DOI: 10.1038/nprot.2012.016
发表时间: 2012-03-01
期刊: Nature protocols
影响因子: 14.8
作者:
通讯作者: --