ALIGNING AMINO-ACID SEQUENCES - COMPARISON OF COMMONLY USED METHODS

ALIGNING AMINO-ACID SEQUENCES - COMPARISON OF COMMONLY USED METHODS
复制标题

DOI:
10.1007/bf02100085
复制
发表时间:
1985-01-01
影响因子:
3.9
通讯作者:
DOOLITTLE, RF
DOOLITTLE, RF
中科院分区:
生物学3区
文献类型:
--
作者:
FENG, DF;JOHNSON, MS;DOOLITTLE, RF

文献摘要

被引文献

相似文献

我们使用四种不同的比对方案检查了两个广泛的蛋白质序列家族,这些方案采用不同程度的“加权”,以确定哪种方法在建立关系时最敏感。所有比对使用基于Needleman和Wunsch设计的通用算法的相似性方法。这些方法包括一个简单的程序UM一种方案是将遗传密码用作加权的基础(GC);另一种方案是采用基于氨基酸结构相似性和突变遗传基础的矩阵(SG);第四种方法使用Dayhoff基于观察到的氨基酸替换开发的经验对数优势矩阵(LOM)。检测的两个序列家族是(a)9种不同的球蛋白和(B)9种不同的酪氨酸激酶样蛋白。人们想当然地认为一个家庭的所有成员都有共同的祖先。在两个序列超过30%相同的情况下,所有四种方法的比对几乎总是相同的。然而,在同一性百分比小于20%的情况下,比对中通常存在显著差异。平均而言,戴霍夫LOM方法在验证远距离关系方面是最有效的,这是由经验性的“混乱测试”判断的。然而,这并不是普遍的情况,在某些情况下,简单的UM实际上同样好,甚至更好。树构建的基础上的各种路线不同的方面,他们的肢体长度,但基本上有相同的分支顺序。我们提出了一些原因,在两个不同的序列设置的四种方法的不同的有效性,并提供一些经验法则来评估序列关系的意义。
We examined two extensive families of protein sequences using four different alignment schemes that employ various degrees of “weighting” in order to determine which approach is most sensitive in establishing relationships. All alignments used a similarity approach based on a general algorithm devised by Needleman and Wunsch. The approaches included a simple program, UM (unitary matrix), whereby only identities are scored; a scheme in which the genetic code is used as a basis for weighting (GC); another that employs a matrix based on structural similarity of amino acids taken together with the genetic basis of mutation (SG); and a fourth that uses the empirical log-odds matrix (LOM) developed by Dayhoff on the basis of observed amino acid replacements. The two sequence families examined were (a) nine different globins and (b) nine different tyrosine kinase-like proteins. It was assumed a priori that all members of a family share common ancestry. In cases where two sequences were more than 30% identical, alignments by all four methods were almost always the same. In cases where the percentage identity was less than 20%, however, there were often significant differences in the alignments. On the average, the Dayhoff LOM approach was the most effective in verifying distant relationships, as judged by an empirical “jumbling test.” This was not universally the case, however, and in some instances the simple UM was actually as good or better. Trees constructed on the basis of the various alignments differed with regard to their limb lengths, but had essentially the same branching orders. We suggest some reasons for the different effectivenesses of the four approaches in the two different sequence settings, and offer some rules of thumb for assessing the significance of sequence relationships.