A MEASURE OF THE SIMILARITY OF SETS OF SEQUENCES NOT REQUIRING SEQUENCE ALIGNMENT
A MEASURE OF THE SIMILARITY OF SETS OF SEQUENCES NOT REQUIRING SEQUENCE ALIGNMENT
复制标题
DOI:
10.1073/pnas.83.14.5155
复制
发表时间:
1986-07-01
影响因子:
11.1
通讯作者:
BLAISDELL, BE
中科院分区:
文献类型:
--
作者:
BLAISDELL, BE
Determination of first- and second-order Markov chain homogeneity of sets of nuclear eukaryotic DNA sequences, both coding and noncoding, finds similarities imperceptible to the standard Needleman-Wunsch base matching or dot-matrix algorithms. These measures of the similarities of the distributions of adjacent pairs or triplets are in agreement with accepted evolutionary-tree topologies. Hierarchical clustering of the distributions of doublets of 30 miscellaneous coding sequences gives clusters in reasonable agreement with accepted biological classification. In addition to similarity by homology, there is also observed similarity of disparate genes in the same organism-for example, all three disparate yeast genes (two enzymes and actin) form a well-distinguished cluster.