LINGUISTIC MEASURE OF TAXONOMIC AND FUNCTIONAL RELATEDNESS OF NUCLEOTIDE-SEQUENCES

LINGUISTIC MEASURE OF TAXONOMIC AND FUNCTIONAL RELATEDNESS OF NUCLEOTIDE-SEQUENCES
复制标题

DOI:
10.1080/07391102.1990.10508563
复制
发表时间:
1990-06-01
影响因子:
4.4
通讯作者:
TRIFONOV, EN
TRIFONOV, EN
中科院分区:
生物学3区
文献类型:
--
作者:
PIETROKOVSKI, S;HIRSHON, J;TRIFONOV, EN

文献摘要

被引文献

相似文献

核苷酸序列中的寡核苷酸“单词”的频率反映了序列“文本”中包含的遗传信息。核苷酸序列特征性地由它们的对比词词汇表表示。序列之间的相关性(不相关性)的比较,通过相关的对比词汇显示,以反映良好的序列之间的相关性。一个单一的值,序列之间的语言相似性,建议作为衡量序列相关性。短至1000个碱基的序列可以通过该技术表征并与其他序列定量相关。语言序列相似性值用于分析分类学上和功能上不同的核苷酸序列。相似性值被证明是非常敏感的源种的相关性,从而提供了一个方便的工具,物种的分类学分类,其序列词汇。功能多样的序列通过它们的语言相似性值而显得不同。这可以是一个快速筛选技术的序列的功能表征和映射功能不同的区域在长序列的基础。
The frequencies of "words", oligonucleotides within nucleotide sequences, reflect the genetic information contained in the sequence "texts". Nucleotide sequences are characteristically represented by their contrast word vocabularies. Comparisons of the sequences by correlating their contrast vocabularies is shown to reflect well the relatedness (unrelatedness) between the sequences. A single value, the linguistic similarity between the sequences, is suggested as a measure of sequences relatedness. Sequences as short as 1000 bases can be characterized and quantitatively related to other sequences by this technique. The linguistic sequence similarity value is used for analysis of taxonomically and functionally diverse nucleotide sequences. The similarity value is shown to be very sensitive to the relatedness of the source species, thus providing a convenient tool for taxonomic classification of species by their sequence vocabularies. Functionally diverse sequences appear distinct by their linguistic similarity values. This can be a basis for a quick screening technique for functional characterization of the sequences and for mapping functionally distinct regions in long sequences.