Alignment by numbers: sequence assembly using compressed numerical representations

Alignment by numbers: sequence assembly using compressed numerical representations
复制标题

按数字对齐:使用压缩数字表示进行序列组装

DOI:
10.1101/011940
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
Tapinos A
Tapinos A
中科院分区:
--
文献类型:
--
作者:
Tapinos A

文献摘要

参考文献

被引文献

相似文献

DNA测序仪器使基因组分析的范围和规模前所未有,扩大了我们生成和解释序列数据的能力之间的差距。已建立的计算序列分析方法通常考虑序列的核苷酸水平分辨率,虽然这些方法足够准确,但日益雄心勃勃的数据密集型分析使它们对于要求苛刻的应用(如基因组和宏基因组组装)来说不切实际。在其他涉及序列数据的数据密集型领域,如信号处理和时间序列分析,也遇到了类似的分析挑战。通过用数字表示核酸组成,有可能将这些领域的降维方法应用于核苷酸序列,使其能够近似表示。为了探索信号分解方法在序列组装中的适用性,我们实现了短读比对器,并评估了其对模拟高多样性病毒序列以及四个现有比对器的性能。使用我们的原型实现,近似序列表示减少了高达14倍的整体比对时间相比,未压缩的序列,并没有任何降低比对精度。尽管使用了高度近似的序列表示,但我们的实现产生了与现有比对器相似的整体准确度的比对,优于在高水平序列变异下测试的所有其他工具。我们的方法也被应用到一个模拟的不同的病毒种群的新组装。我们已经证明,全序列分辨率不是准确序列比对的先决条件,并且通过适当的序列降维可以保留甚至增强分析性能。
DNA sequencing instruments are enabling genomic analyses of unprecedented scope and scale, widening the gap between our abilities to generate and interpret sequence data. Established methods for computational sequence analysis generally consider the nucleotide-level resolution of sequences, and while these approaches are sufficiently accurate, increasingly ambitious and data-intensive analyses are rendering them impractical for demanding applications such as genome and metagenome assembly. Comparable analytical challenges are encountered in other data-intensive fields involving sequential data such as signal processing and time series analysis. By representing nucleic acid composition numerically it is possible to apply dimensionality reduction methods from these fields to sequences of nucleotides, enabling their approximate representation. To explore the applicability of signal decomposition methods in sequence assembly, we implemented a short read aligner and evaluated its performance against simulated high diversity viral sequences alongside four existing aligners. Using our prototype implementation, approximate sequence representations reduced overall alignment time by up to 14-fold compared to that of uncompressed sequences, and without any reduction in alignment accuracy. Despite using heavily approximated sequence representations, our implementation yielded alignments of similar overall accuracy to existing aligners, outperforming all other tools tested at high levels of sequence variation. Our approach was also applied to thede novoassembly of a simulated diverse viral population. We have demonstrated that full sequence resolution is not a prerequisite of accurate sequence alignment and that analytical performance may be retained or even enhanced through appropriate dimensionality reduction of sequences.
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1101/gr.101360.109
发表时间: 2010-09-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Schatz, Michael C.;Delcher, Arthur L.;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1038/nature03959
发表时间: 2005-09-15
期刊: NATURE
影响因子: 64.8
作者:
Margulies, M;Egholm, M;Rothberg, JM
通讯作者: Rothberg, JM
DOI: 10.1093/bioinformatics/btu146
发表时间: 2014-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Otto, Christian;Stadler, Peter F.;Hoffmann, Steve
通讯作者: Hoffmann, Steve
DOI: 10.1101/gr.126599.111
发表时间: 2011-12-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Earl, Dent;Bradnam, Keith;Paten, Benedict
通讯作者: Paten, Benedict