A MEASURE OF THE SIMILARITY OF SETS OF SEQUENCES NOT REQUIRING SEQUENCE ALIGNMENT

A MEASURE OF THE SIMILARITY OF SETS OF SEQUENCES NOT REQUIRING SEQUENCE ALIGNMENT
复制标题

DOI:
10.1073/pnas.83.14.5155
复制
发表时间:
1986-07-01
影响因子:
11.1
通讯作者:
BLAISDELL, BE
BLAISDELL, BE
中科院分区:
综合性期刊1区
文献类型:
--
作者:
BLAISDELL, BE

文献摘要

被引文献

相似文献

一阶和二阶马尔可夫链同质性的核真核生物DNA序列集的确定,编码和非编码,发现难以察觉的标准Needleman-Wunsch碱基匹配或点阵算法的相似性。这些措施的相似性分布的相邻对或三胞胎是在协议接受进化树拓扑结构。30个混杂编码序列的双联体分布的分层聚类给出与公认的生物分类合理一致的聚类。除了同源性的相似性之外,在同一生物体中还观察到不同基因的相似性,例如,所有三个不同的酵母基因(两种酶和肌动蛋白)形成了一个非常明显的簇。
Determination of first- and second-order Markov chain homogeneity of sets of nuclear eukaryotic DNA sequences, both coding and noncoding, finds similarities imperceptible to the standard Needleman-Wunsch base matching or dot-matrix algorithms. These measures of the similarities of the distributions of adjacent pairs or triplets are in agreement with accepted evolutionary-tree topologies. Hierarchical clustering of the distributions of doublets of 30 miscellaneous coding sequences gives clusters in reasonable agreement with accepted biological classification. In addition to similarity by homology, there is also observed similarity of disparate genes in the same organism-for example, all three disparate yeast genes (two enzymes and actin) form a well-distinguished cluster.