Multivariate entropy distance method for prokaryotic gene identification.

Multivariate entropy distance method for prokaryotic gene identification.
复制标题

DOI:
10.1142/s0219720004000624
复制
发表时间:
2004-06-01
影响因子:
1
通讯作者:
She, Zhen-Su
She, Zhen-Su
中科院分区:
生物学4区
文献类型:
--
作者:
Ouyang, Zhengqing;Zhu, Huaiqiu;She, Zhen-Su

文献摘要

被引文献

相似文献

为原核生物基因组编码序列的准确、高效鉴定提供了一种新的简便方法。该方法采用香农描述DNA序列的人工语言。它包括根据通用遗传密码将DNA序列翻译成包含20个基本单词的伪氨基酸序列。该方法利用熵密度剖面(EDP),将有限长度的序列映射到一个矢量上,然后根据其性质分析其在20维相空间中的位置。研究发现,在少量(最多1个)开放阅读帧(orf)上,编码和非编码平均EDP的相对距离之比可以作为一个较好的编码电位。设计了一种迭代算法,利用这种编码势来寻找一组“根”序列。然后提出了一种多变量熵距离(MED)算法用于原核生物基因的识别;它的特点是结合使用编码电位和基于edp的序列相似性分析。当前版本的MED是无监督的,无参数的,易于实现。在NCBI的RefSeq数据库中,检测结果表明,该方法能够检测出95-99%的基因和10-30%的附加基因,检测出97.5-99.8%的已知功能的已确认基因。它还被证明能够找到一组(功能已知的)基因,这些基因被其他知名的基因寻找算法所遗漏。实验结果表明,MED算法在原核生物基因预测方面达到了与GeneMark和Glimmer等算法相似的性能水平。
A new simple method is found for efficient and accurate identification of coding sequences in prokaryotic genome. The method employs a Shannon description of artificial language for DNA sequences. It consists in translating a DNA sequence into a pseudo-amino acid sequence with 20 fundamental words according to the universal genetic code. With an entropy-density profile (EDP), the method maps a sequence of finite length to a vector and then analyzes its position in the 20-dimensional phase space depending on its nature. It is found that the ratio of the relative distance to an averaged coding and non-coding EDP over a small number (up to one) of open reading frames (ORFs) can serve as a good coding potential. An iterative algorithm is designed for finding a set of "root" sequences using this coding potential. A multivariate entropy distance (MED) algorithm is then proposed for the identification of prokaryotic genes; it has a feature to combine the use of a coding potential and an EDP-based sequence similarity analysis. The current version of MED is unsupervised, parameter-free and simple to implement. It is demonstrated to be able to detect 95-99% genes with 10-30% of additional genes when tested against the RefSeq database of NCBI and to detect 97.5-99.8% of confirmed genes with known functions. It is also shown to be able to find a set of (functionally known) genes that are missed by other well-known gene finding algorithms. All measurements show that the MED algorithm reaches a similar performance level as the algorithms like GeneMark and Glimmer for prokaryotic gene prediction.