Prediction of complete gene structures in human genomic DNA

Prediction of complete gene structures in human genomic DNA
复制标题

DOI:
10.1006/jmbi.1997.0951
复制
发表时间:
1997-04-25
影响因子:
5.6
通讯作者:
Karlin, S
Karlin, S
中科院分区:
生物学2区
文献类型:
--
作者:
Burge, C;Karlin, S

文献摘要

被引文献

相似文献

我们介绍了一个通用的概率模型的基因结构的人类基因组序列,其中包括基本的转录,翻译和剪接信号的描述,以及。外显子、内含子和基因间区的长度分布和组成特征。不同的模型参数集的推导,以解释在不同的C + G组成区域的人类基因组中观察到的基因密度和结构的许多实质性差异。此外,新的模型的供体和受体剪接信号的描述,捕获信号位置之间的潜在重要的依赖关系。该模型被应用到一个计算机程序,GENSCAN,它确定完整的外显子/内含子结构的基因组DNA中的基因的基因识别问题。该程序的新功能包括预测序列中多个基因的能力,处理部分和完整的基因,以及预测发生在一条或两条DNA链上的一致基因组。GENSCAN被证明具有比现有方法更高的准确性,当测试标准化的人类和脊椎动物基因集时,准确识别出75%至80%的外显子。该程序还能够相当准确地指示每个预测外显子的可靠性。对于不同C + G含量的序列和不同的脊椎动物组,观察到一致的lv高水平的准确性。(C)出版社:Academic Press Limited。
We introduce a general probabilistic model of the gene structure of human genomic sequences which incorporates descriptions of the basic transcriptional, translational and splicing signals, as well. as length distributions and compositional features of exons, introns and intergenic regions. Distinct sets of model parameters are derived to account for the many substantial differences in gene density and structure observed in distinct C + G compositional regions of the human genome. Lu addition, new models of the donor and acceptor splice signals are described which capture potentially important dependencies between signal positions. The model is applied to the problem of gene identification in a computer program, GENSCAN, which identifies complete exon/intron structures of genes in genomic DNA. Novel features of the program include the capacity to predict multiple genes in a sequence, to deal with partial as well as complete genes, and to predict consistent sets of genes occurring on either or both DNA strands. GENSCAN is shown to have substantially higher accuracy than existing methods when tested on standardized sets of human and vertebrate genes, with 75 to 80% of exons identified exactly. The program is also capable of indicating fairly accurately the reliability of each predicted exon. Consistent lv high levels of accuracy are observed for sequences of differing C + G content and for distinct groups of vertebrates. (C) 1997 Academic Press Limited.