LINGUISTIC MEASURE OF TAXONOMIC AND FUNCTIONAL RELATEDNESS OF NUCLEOTIDE-SEQUENCES
LINGUISTIC MEASURE OF TAXONOMIC AND FUNCTIONAL RELATEDNESS OF NUCLEOTIDE-SEQUENCES
复制标题
DOI:
10.1080/07391102.1990.10508563
复制
发表时间:
1990-06-01
影响因子:
4.4
通讯作者:
TRIFONOV, EN
中科院分区:
文献类型:
--
作者:
PIETROKOVSKI, S;HIRSHON, J;TRIFONOV, EN
The frequencies of "words", oligonucleotides within nucleotide sequences, reflect the genetic information contained in the sequence "texts". Nucleotide sequences are characteristically represented by their contrast word vocabularies. Comparisons of the sequences by correlating their contrast vocabularies is shown to reflect well the relatedness (unrelatedness) between the sequences. A single value, the linguistic similarity between the sequences, is suggested as a measure of sequences relatedness. Sequences as short as 1000 bases can be characterized and quantitatively related to other sequences by this technique. The linguistic sequence similarity value is used for analysis of taxonomically and functionally diverse nucleotide sequences. The similarity value is shown to be very sensitive to the relatedness of the source species, thus providing a convenient tool for taxonomic classification of species by their sequence vocabularies. Functionally diverse sequences appear distinct by their linguistic similarity values. This can be a basis for a quick screening technique for functional characterization of the sequences and for mapping functionally distinct regions in long sequences.