MAVID: Constrained ancestral alignment of multiple sequences

MAVID: Constrained ancestral alignment of multiple sequences
复制标题

DOI:
10.1101/gr.1960404
复制
发表时间:
2004-04-01
期刊:
影响因子:
7
通讯作者:
Pachter, L
Pachter, L
中科院分区:
生物学1区
文献类型:
--
作者:
Bray, N;Pachter, L

文献摘要

被引文献

相似文献

我们描述了一个新的全球多重比对程序能够对齐大量的基因组区域。我们的渐进比对方法采用了以下思路:祖先序列的最大似然推断,自动引导树的构建,基于蛋白质的锚定从头基因预测,和来自全球同源性序列图的约束。我们已经在MAVID程序中实现了这些想法,该程序能够准确地将多个基因组区域LIP对齐到兆字节长。MAVID能够有效地比对不同的序列,以及不完整的未完成的序列。我们展示了该计划的基准CFTR区域,其中包括1.8 Mb的人类序列和20 orthopathic有袋动物,鸟类,鱼类和哺乳动物的区域的能力。最后,我们描述了两个大的MAVID比对,一个是所有可用的HIV基因组的比对,另一个是整个人类、小鼠和大鼠基因组的多重比对。
We describe a new global multiple-alignment program capable of aligning a large number of genomic regions. Our progressive-alignment approach incorporates the following ideas: maximum-likelihood inference of ancestral sequences, automatic guide-tree construction, protein-based anchoring of ab-initio gene predictions, and constraints derived from a global homology map of the sequences. We have implemented these ideas in the MAVID program, which is able to accurately align multiple genomic regions LIP to megabases long. MAVID is able to effectively align divergent sequences, as well as incomplete unfinished sequences. We demonstrate the capabilities of the program on the benchmark CFTR region, which consists of 1.8 Mb of human sequence and 20 orthologous regions in marsupials, birds, fish, and mammals. Finally, we describe two large MAVID alignments, an alignment of all the available HIV genomes and a multiple alignment of the entire human, mouse, and rat genomes.