MULTIPLE SEQUENCE ALIGNMENT

MULTIPLE SEQUENCE ALIGNMENT
复制标题

DOI:
10.1016/0022-2836(86)90252-4
复制
发表时间:
1986-09-20
影响因子:
5.6
通讯作者:
ANDERSON, WF
ANDERSON, WF
中科院分区:
生物学2区
文献类型:
--
作者:
BACON, DJ;ANDERSON, WF

文献摘要

被引文献

相似文献

已经开发出一种用于同时比对几个序列的片段的方法。搜索步数只取决于序列的数量,而不是指数,因为大多数比对在没有显式评估的情况下被拒绝。本文中称为“堆”的数据结构促进了这一过程。对于一组n个序列片段,总相似度取为所有组成片段对相似度的和,其又是来自表格的相应残基相似度分数的和。测试比对重要性的统计模型使客观地对序列进行分组成为可能,即使大多数或所有的相互关系都很弱。这些测试非常敏感,同时保持相当保守,并不鼓励在现有的集合中添加“不匹配”的序列。新技术被应用于一组五种DNA结合蛋白,一组使用辅酶FAD的三种酶,以及一组对照。先前提出的基于结构比较和序列检查的DNA结合蛋白的比对得到了相当大的支持,并且发现了FAD结合蛋白的高度显著的比对。
A method has been developed for aligning segments of several sequences at once. The number of search steps depends only polynomially on the number of sequences, instead of exponentially, because most alignments are rejected without being evaluated explicitly. A data structure herein called the "heap" facilitates this process. For a set of n sequence segments, the overall similarity is taken to the sum of all the constituent segment pair similariteis, which are in turn sums of corresponding residue similarity scores from a Table. The statistical models that test alignments for significance make it possible to group sequences objectively, even when most or all of the interrelationships are weak. These tests are very sensitive, while remaining quite conservative, and discourage the addition of "misfit" sequences to an existing set. The new techniques are applied to a set of five DNA-binding proteins, to a group of three enzymes that employ the coenzyme FAD, and to a control set. The alignment previously proposed for the DNA-binding proteins on the basis of structural comparisons and inspection of sequences is supported quite dramatically, and a highly significant alignment is found for the FAD-binding proteins.