Recognizing exons in genomic sequence using GRAIL II.

Recognizing exons in genomic sequence using GRAIL II.
复制标题

使用 GRAIL II 识别基因组序列中的外显子。

DOI:
--
复制
发表时间:
1994
期刊:
Genetic engineering
影响因子:
--
通讯作者:
E. Uberbacher
E. Uberbacher
中科院分区:
--
文献类型:
--
作者:
Ying Xu;R. Mural;Manesh B Shah;E. Uberbacher

文献摘要

被引文献

相似文献

我们描述了一种改进的神经网络系统,用于识别人类基因组DNA序列中的蛋白质编码区(外显子)。该编码区识别系统是GRAIL的新版本GRAIL II的一部分,比以前的GRAIL系统的编码识别性能有了显著的改进。Grail II将定位外显子的过程分为四个步骤。它首先产生一个外显子候选库,由测试序列的所有开放阅读框内的所有可能的(翻译起始-供体)、(受体-供体)和(受体-翻译停止)对组成。通过应用一套启发式规则,这些外显子候选中的绝大多数都被排除在考虑之外。在减少候选库的大小后,GRAIL II使用三个训练好的神经网络来评估起始外显子、内部外显子和末端外显子候选边缘的编码潜力和准确性。这些网络为每个外显子输出一组重叠的候选基因,这些外显子的分数和边缘位置不同。给定外显子的多个候选基于它们相对于对应于其他外显子的候选的位置被分组为一簇,并且每个簇的得分最高的候选被用作相应外显子的“最佳”预测。与以前的GRAIL版本不同,GRAIL II使用可变长度的窗口来评估外显子候选,其性能几乎与外显子长度无关。除了编码潜力的几个强有力的指标外,该系统还使用其他几种类型的信息,包括剪接连接的分数、GC组成以及外显子候选相邻区域的属性,以帮助识别过程。在Genbank(3)的一大组序列中,GRAIL II定位了93%的所有外显子,而不考虑大小,假阳性率为12%。在真正的阳性中,62%与实际外显子完全匹配(外显子边与碱基正确),93%至少与一条边正确匹配。通过GRAIL II的基因组装程序(GAP III)(4)模块构建基因模型的过程,进一步提高了这些统计量,特别是边缘的假阳性率和准确性,该模块使用评分的外显子候选作为输入,构建最佳的基因模型。基因建模系统将在其他地方介绍。
We have described an improved neural network system for recognizing protein coding regions (exons) in human genomic DNA sequences. This coding region recognition system is part of a new version of GRAIL, GRAIL II, and represents a significant improvement over the coding recognition performance of the previous GRAIL system. GRAIL II divides the process of locating exons into four steps. It first generates an exon candidate pool consisting of all possible (translation start-donor), (acceptor-donor), and (acceptor-translation stop) pairs within all open reading frames of the test sequence. The vast majority of these exon candidates are eliminated from consideration by applying a set of heuristic rules. After reducing the size of the candidate pool, GRAIL II uses three trained neural networks to evaluate the coding potential and accuracy of the edges of starting exon, internal exon and terminal exon candidates. These networks output a set of overlapping candidates for each exon which differ by their scores and position of their edges. Multiple candidates for a given exon are grouped into a cluster based on their locations relative to candidates corresponding to other exons, and the highest scoring candidate for each cluster is used as the "best" prediction of the corresponding exon. Unlike the previous GRAIL version, GRAIL II uses variable-length windows to evaluate exon candidates and its performance is nearly independent of exon length. In addition to several strong indicators of coding potential, the system uses several other types of information including scores for splice junctions, GC composition, and the properties of the regions adjacent to an exon candidate, to aid in the discrimination process. On a large set of sequences from Genbank (3), GRAIL II located 93% of all exons regardless of size with a false positive rate of 12%. Among the true positives, 62% match the actual exons exactly (the exons edges are correct to the base), and 93% match at least one edge correctly. These statistics are further improved, especially the false positive rate and accuracy of the edges, through a process of gene model construction by the Gene Assembly Program (GAP III) (4) module of GRAIL II, which uses the scored exon candidates as input and constructs optimal gene models. The gene modeling system will be described elsewhere.