JIGSAW, GeneZilla, and GlimmerHMM: puzzling out the features of human genes in the ENCODE regions.

JIGSAW, GeneZilla, and GlimmerHMM: puzzling out the features of human genes in the ENCODE regions.
复制标题

DOI:
10.1186/gb-2006-7-s1-s9
复制
发表时间:
2006
期刊:
影响因子:
12.3
通讯作者:
Salzberg SL
Salzberg SL
中科院分区:
生物学1区
文献类型:
--
作者:
Allen JE;Majoros WH;Pertea M;Salzberg SL

文献摘要

参考文献

被引文献

相似文献

预测人类DNA中完整的蛋白质编码基因仍然是一个重大挑战。虽然已经研究了许多有前途的方法,但尚未出现一套理想的工具,可以在整个基因水平上提供近乎完美的灵敏度和特异性水平。作为在这个方向上的一个渐进步骤,希望在ENCODE区域的受控基因发现实验将提供一个更准确的观点,不同的策略建模和预测基因结构的相对优势。在这里,我们描述了我们的通用真核生物基因发现管道及其主要组成部分,以及我们发现在我们的管道中容纳人类DNA所必需的方法学适应,注意到当一个新的基因组类别提交给社区进行分析时,我们和其他具有类似管道的人可能需要类似水平的努力。我们还描述了一些受控实验,涉及到不同类型的证据和特征状态的差异纳入我们的模型,以及这些变化对预测准确性的影响。虽然在非比较基因发现者的情况下,我们发现添加模型状态来表示特定的生物特征对提高预测准确性几乎没有作用,但对于我们的基于证据的“组合器”程序,合并额外的证据跟踪往往会产生显着的收益对于大多数证据类型,这表明在隐马尔可夫模型水平上改进建模工作的价值相对较小。我们将这些发现与我们当前的未来研究计划联系起来。
Predicting complete protein-coding genes in human DNA remains a significant challenge. Though a number of promising approaches have been investigated, an ideal suite of tools has yet to emerge that can provide near perfect levels of sensitivity and specificity at the level of whole genes. As an incremental step in this direction, it is hoped that controlled gene finding experiments in the ENCODE regions will provide a more accurate view of the relative benefits of different strategies for modeling and predicting gene structures. Here we describe our general-purpose eukaryotic gene finding pipeline and its major components, as well as the methodological adaptations that we found necessary in accommodating human DNA in our pipeline, noting that a similar level of effort may be necessary by ourselves and others with similar pipelines whenever a new class of genomes is presented to the community for analysis. We also describe a number of controlled experiments involving the differential inclusion of various types of evidence and feature states into our models and the resulting impact these variations have had on predictive accuracy. While in the case of the non-comparative gene finders we found that adding model states to represent specific biological features did little to enhance predictive accuracy, for our evidence-based 'combiner' program the incorporation of additional evidence tracks tended to produce significant gains in accuracy for most evidence types, suggesting that improved modeling efforts at the hidden Markov model level are of relatively little value. We relate these findings to our current plans for future research.
DOI: 10.1186/1471-2105-6-16
发表时间: 2005-01-24
期刊: BMC bioinformatics
影响因子: 3
作者:
Majoros WH;Pertea M;Delcher AL;Salzberg SL
通讯作者: Salzberg SL
DOI: 10.1093/bioinformatics/btg1080
发表时间: 2003-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Stanke, Mario;Waack, Stephan
通讯作者: Waack, Stephan
DOI: 10.1093/nar/gkg033
发表时间: 2003-01-01
影响因子: 14.9
作者:
Wheeler, DL;Church, DM;Wagner, L
通讯作者: Wagner, L
DOI: 10.1101/gr.1858004
发表时间: 2004-05-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Curwen, V;Eyras, E;Clamp, M
通讯作者: Clamp, M
DOI: 10.1093/bioinformatics/19.2.219
发表时间: 2003-01-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Pedersen, JS;Hein, J
通讯作者: Hein, J