ASSESSMENT OF PROTEIN CODING MEASURES

ASSESSMENT OF PROTEIN CODING MEASURES
复制标题

DOI:
10.1093/nar/20.24.6441
复制
发表时间:
1992-12-25
影响因子:
14.9
通讯作者:
TUNG, CS
TUNG, CS
中科院分区:
生物学2区
文献类型:
--
作者:
FICKETT, JW;TUNG, CS

文献摘要

被引文献

相似文献

在过去的13年中,已经发表了许多用于识别DNA序列中的蛋白质编码基因的方法,并且继续开发利用现有技术的新的、更全面的算法。为了优化持续发展,系统地回顾和评估已发表的技术是有价值的。大多数基因识别算法的核心是一个或多个编码度量-给定序列的任何样本窗口,产生用于测量样本序列与“典型”外显子DNA窗口相似程度的数字或向量的函数。在本文中,我们回顾和合成的基本编码措施,从出版的算法。描述了一个标准化的基准,并根据该基准对每个措施进行评估。我们的主要结论是,一个非常简单和明显的措施-计数寡聚体-比任何更复杂的措施更有效。不同的测量包含不同的信息。然而,目前的一套措施存在大量冗余。我们表明,在基因识别算法的未来发展中,注意力可能会被限制到6的20个左右的措施提出的日期。
A number of methods for recognizing protein coding genes in DNA sequence have been published over the last 13 years, and new, more comprehensive algorithms, drawing on the repertoire of existing techniques, continue to be developed. To optimize continued development, it is valuable to systematically review and evaluate published techniques. At the core of most gene recognition algorithms is one or more coding measures - functions which produce, given any sample window of sequence, a number or vector intended to measure the degree to which a sample sequence resembles a window of 'typical' exonic DNA. In this paper we review and synthesize the underlying coding measures from published algorithms. A standardized benchmark is described, and each of the measures is evaluated according to this benchmark. Our main conclusion is that a very simple and obvious measure - counting oligomers - is more effective than any of the more sophisticated measures. Different measures contain different information. However there is a great deal of redundancy in the current suite of measures. We show that in future development of gene recognition algorithms, attention can probably be limited to six of the twenty or so measures proposed to date.