EFFECTIVENESS OF MEASURES REQUIRING AND NOT REQUIRING PRIOR SEQUENCE ALIGNMENT FOR ESTIMATING THE DISSIMILARITY OF NATURAL SEQUENCES

EFFECTIVENESS OF MEASURES REQUIRING AND NOT REQUIRING PRIOR SEQUENCE ALIGNMENT FOR ESTIMATING THE DISSIMILARITY OF NATURAL SEQUENCES
复制标题

DOI:
10.1007/bf02602924
复制
发表时间:
1989-01-01
影响因子:
3.9
通讯作者:
BLAISDELL, BE
BLAISDELL, BE
中科院分区:
生物学3区
文献类型:
--
作者:
BLAISDELL, BE

文献摘要

被引文献

相似文献

通过无根进化树的边(分支长度)的加性最小二乘估计与观察到的两两不相似度量的拟合程度以及从同一序列集派生的不同数据集的树的一致性来评估序列不相似的各种度量。该评价在不同度量和可能树之间提供了敏感的区分。不需要先验序列比对的不相似度测量方法与需要先验序列比对的传统错配计数方法一样好。应用Jukes-Cantor校正单线态失配计数使结果恶化。不需要校准的测量方法具有适用于差异太大而不能严格校准的序列的优点。已经使用了两种不同的不需要比对的成对不相似性度量方法:(1)多重分布距离(MDD),即各自序列中碱基单片段(或双片段,或三元组,或…)的部分向量之间的欧几里得距离的平方,以及(2)长词补(CLW),即不出现在显著长常用词中的碱基计数。MDD适用于比CLW(非编码)更不同的序列,但是后者通常在两种度量都可用(编码)的情况下给出更好的结果。通过使用更长的多胞胎,如果序列编码,使用更大的氨基酸和密码子字母表而不是核苷酸字母表,MDD结果得到改善。加性最小二乘法可以为同一物种(或相关基因)的不同树提供合理的一致性。
Various measures of sequence dissimilarity have been evaluated by how well the additive least squares estimation of edges (branch lengths) of an unrooted evolutionary tree fit the observed pairwise dissimilarity measures and by how consistent the trees are for different data sets derived from the same set of sequences. This evaluation provided sensitive discrimination among dissimilarity measures and among possible trees. Dissimilarity measures not requiring prior sequence alignment did about as well as did the traditional mismatch counts requiring prior sequence alignment. Application of Jukes-Cantor correction to singlet mismatch counts worsened the results. Measures not requiring alignment had the advantage of being applicable to sequences too different to be critically alignable. Two different measures of pairwise dissimilarity not requiring alignment have been used: (1) multiplet distribution distance (MDD), the square of the Euclidean distance between vectors of the fractions of base singlets (or doublets, or triplets, or ...) in the respective sequences, and (2) complements of long words (CLW), the count of bases not occurring in significantly long common words. MDD was applicable to sequences more different than was CLW (noncoding), but the latter often gave better results where both measures were available (coding). MDD results were improved by using longer multiplets and, if the sequences were coding, by using the larger amino acid and codon alphabets rather than the nucleotide alphabet. The additive least squares method could be used to provide a reasonable consensus of different trees for the same set of species (or related genes).