Fast statistical alignment.

Fast statistical alignment.
复制标题

DOI:
10.1371/journal.pcbi.1000392
复制
发表时间:
2009-05
影响因子:
4.3
通讯作者:
Pachter L
Pachter L
中科院分区:
生物学2区
文献类型:
--
作者:
Bradley RK;Roberts A;Smoot M;Juvekar S;Do J;Dewey C;Holmes I;Pachter L

文献摘要

参考文献

被引文献

相似文献

我们描述了一个新的程序,用于多个生物序列的比对,这是统计动机和足够快的问题大小,在实践中出现的。我们的快速统计比对程序是基于对隐马尔可夫模型,它近似的插入/删除过程的树,并使用序列退火算法联合收割机的后验概率估计从这些模型到一个多个对齐。FSA使用其明确的统计模型,以产生多个对齐,这是伴随着估计的对齐精度和不确定性的每列和字符的比对-以前只有与对齐程序,使用计算昂贵的马尔可夫链蒙特卡罗方法-但可以对齐数千个长序列。此外,FSA利用一个无监督的查询特定的学习过程的参数估计,从而提高基准参考比对的准确性相比,现有的程序。与其他方法相比,FSA采用的质心对齐方法与其学习过程相结合,大大减少了生物数据上的假阳性对齐量。FSA程序和用于探索比对中的不确定性的伴随可视化工具可以经由http://orangutan.math.berkeley.edu/fsa/的网络界面使用,并且源代码可以在http://fsa.sourceforge.net/获得。生物序列比对是比较基因组学的基本问题之一,至今尚未得到解决。维基百科上列出了60多个序列比对程序,每年都有许多新程序发布。然而,许多流行的程序遭受病理,如对齐不相关的序列,并在蛋白质(氨基酸)和密码子(核苷酸)空间产生不一致的对齐,怀疑推断的对齐的准确性。不准确的比对会给下游分析(如系统发育树重建和替代率估计)带来巨大且未知的系统偏倚。我们描述了一个新的程序,多序列比对,可以对齐蛋白质,RNA和DNA序列,并提高了现有的方法对蛋白质和RNA的结构比对和模拟哺乳动物和苍蝇的基因组比对的基准的准确性。我们的方法,它试图找到的比对,这是最接近我们的统计模型下的真相,留下不相关的序列在很大程度上未对齐,并产生一致的比对蛋白质和密码子空间。它对于一些困难的问题来说是足够快的,例如对齐正向同源基因组区域或对齐数百或数千个蛋白质。此外,它还有一个配套的GUI,用于可视化估计的对准可靠性。
We describe a new program for the alignment of multiple biological sequences that is both statistically motivated and fast enough for problem sizes that arise in practice. Our Fast Statistical Alignment program is based on pair hidden Markov models which approximate an insertion/deletion process on a tree and uses a sequence annealing algorithm to combine the posterior probabilities estimated from these models into a multiple alignment. FSA uses its explicit statistical model to produce multiple alignments which are accompanied by estimates of the alignment accuracy and uncertainty for every column and character of the alignment—previously available only with alignment programs which use computationally-expensive Markov Chain Monte Carlo approaches—yet can align thousands of long sequences. Moreover, FSA utilizes an unsupervised query-specific learning procedure for parameter estimation which leads to improved accuracy on benchmark reference alignments in comparison to existing programs. The centroid alignment approach taken by FSA, in combination with its learning procedure, drastically reduces the amount of false-positive alignment on biological data in comparison to that given by other methods. The FSA program and a companion visualization tool for exploring uncertainty in alignments can be used via a web interface at http://orangutan.math.berkeley.edu/fsa/, and the source code is available at http://fsa.sourceforge.net/. Biological sequence alignment is one of the fundamental problems in comparative genomics, yet it remains unsolved. Over sixty sequence alignment programs are listed on Wikipedia, and many new programs are published every year. However, many popular programs suffer from pathologies such as aligning unrelated sequences and producing discordant alignments in protein (amino acid) and codon (nucleotide) space, casting doubt on the accuracy of the inferred alignments. Inaccurate alignments can introduce large and unknown systematic biases into downstream analyses such as phylogenetic tree reconstruction and substitution rate estimation. We describe a new program for multiple sequence alignment which can align protein, RNA and DNA sequence and improves on the accuracy of existing approaches on benchmarks of protein and RNA structural alignments and simulated mammalian and fly genomic alignments. Our approach, which seeks to find the alignment which is closest to the truth under our statistical model, leaves unrelated sequences largely unaligned and produces concordant alignments in protein and codon space. It is fast enough for difficult problems such as aligning orthologous genomic regions or aligning hundreds or thousands of proteins. It furthermore has a companion GUI for visualizing the estimated alignment reliability.
DOI: 10.1038/nature06341
发表时间: 2007-11-08
期刊: NATURE
影响因子: 64.8
作者:
Clark, Andrew G.;Eisen, Michael B.;MacCallum, Iain
通讯作者: MacCallum, Iain
DOI: 10.1093/bioinformatics/bti1200
发表时间: 2005-01-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cartwright, RA
通讯作者: Cartwright, RA
DOI: 10.1186/1471-2105-7-400
发表时间: 2006-09-04
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
D Dowell, Robin;Eddy, Sean R.
通讯作者: Eddy, Sean R.
DOI: 10.1101/gr.1960404
发表时间: 2004-04-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Bray, N;Pachter, L
通讯作者: Pachter, L
DOI: 10.1093/nar/gkh361
发表时间: 2004-07-01
影响因子: 14.9
作者:
Brudno, M;Steinkamp, R;Morgenstern, B
通讯作者: Morgenstern, B