PREDICTION OF GENE STRUCTURE

PREDICTION OF GENE STRUCTURE
复制标题

DOI:
10.1016/0022-2836(92)90130-c
复制
发表时间:
1992-07-05
影响因子:
5.6
通讯作者:
SMITH, T
SMITH, T
中科院分区:
生物学2区
文献类型:
--
作者:
GUIGO, R;KNUDSEN, S;SMITH, T

文献摘要

被引文献

相似文献

我们已经开发了一个层次化的规则库系统,用于识别DNA序列中的基因。原子位点(如起始密码子、终止密码子、受体位点和供体位点)通过许多不同的方法鉴定,并通过一组筛选器和选择的规则进行评估,以最大限度地提高灵敏度;这些被组合成更高阶的基因元素(如外显子),进行评估、筛选并作为等价类组合成可能的基因,对其进行评估和排名。该系统已经在小于15,000个碱基的脊椎动物基因的广泛收集上进行了测试。结果表明,平均而言,88%的预测编码区的转录单位是实际编码,80%的实际编码是正确预测。在大多数应用中,这将足以对蛋白质序列数据库进行搜索,以识别可能的基因功能。此外,该系统提供了一个通用的测试平台,既基因原子位点识别和规则的评估和组装。
We have developed a hierarchical rule base system for identifying genes in DNA sequences. Atomic sites (such as initiation codons, stop codons, acceptor sites and donor sites) are identified by a number of different methods and evaluated by a set of filters and rules chosen to maximize sensitivity; these are combined into higher-order gene elements (such as exons), evaluated, filtered and combined as equivalence classes into probable genes, which are evaluated and ranked. The system has been tested on an extensive collection of vertebrate genes smaller than 15,000 bases. Results obtained show that, on average, 88% of the predicted coding region for a transcription unit is actually coding, and 80% of the actual coding is correctly predicted. This will, in most applications, be sufficient for a search against protein sequence databases for the identification of probable gene function. In addition, the system provides a general test platform for both gene atomic site identification and the rules for their evaluation and assembly.