An assessment of gene prediction accuracy in large DNA sequences

An assessment of gene prediction accuracy in large DNA sequences
复制标题

DOI:
10.1101/gr.122800
复制
发表时间:
2000-10-01
期刊:
影响因子:
7
通讯作者:
Fickett, JW
Fickett, JW
中科院分区:
生物学1区
文献类型:
--
作者:
Guigó, R;Agarwal, P;Fickett, JW

文献摘要

被引文献

相似文献

人类基因组的首批有用产品之一将是一组可预测的基因。除了其内在的科学兴趣外,该数据集的准确性和完整性对人类健康和医学具有相当重要的意义。尽管在计算基因鉴定的方法和准确性评估措施方面已经取得了进展,但大多数用于测试程序的序列集都是短基因组序列,并且人们担心这些准确性措施可能无法很好地外推到更大,更具挑战性的数据集。鉴于缺乏经过实验验证的大型基因组数据集,我们构建了一个半人工测试集,该测试集包括许多短的单基因基因组序列,随机生成基因间区域。这个测试集,应该仍然比真正的人类基因组序列更容易出现一个问题,它模拟了类似于200kb长的BACs被测序。在我们对这些较长的基因组序列的实验中,最准确的从头计算基因预测程序之一GENSCAN的准确性显着下降,尽管它的灵敏度仍然很高。相反,基于相似性的程序(如GENEWISE、PROCRUSTES和BLASTX)的准确性不受随机基因间序列存在的显著影响,而是取决于与蛋白质同源物的相似性强度。正如预期的那样,如果使用更遥远的同系物构建模型,则准确性会下降,并且我们能够定量地估计这种下降。然而,即使在相似性较弱的情况下,这些技术的特异性仍然很好,这是驱动昂贵的后续实验的理想特性。我们的实验表明,虽然基因预测将随着每一种新蛋白质的发现和现有工具的改进而改进,但在我们用纯粹的计算方法破译人类基因组中每个基因的精确外显子结构之前,我们还有很长的路要走。
One of the first useful products From the human genome will be a set of predicted genes. Besides its intrinsic scientific interest, the accuracy and completeness of this data set is of considerable importance for human health and medicine. Though progress has been made on computational gene identification in terms of both methods and accuracy evaluation measures, most of the sequence sets in which the programs are tested are short genomic sequences, and there is concern that these accuracy measures may not extrapolate well to larger, more challenging data sets. Given the absence of experimentally verified large genomic data sets, we constructed a semiartificial test set comprising a number of short single-gene genomic sequences with randomly generated intergenic regions. This test set, which should still present an easier problem than real human genomic sequence, mimics the similar to 200kb long BACs being sequenced. In our experiments with these longer genomic sequences, the accuracy of GENSCAN, one of the most accurate ab initio gene prediction programs, dropped significantly, although its sensitivity remained high. Conversely, the accuracy of similarity-based programs, such as GENEWISE, PROCRUSTES, and BLASTX, was not affected significantly by the presence of random intergenic sequence, but depended on the strength of the similarity to the protein homolog. As expected, the accuracy dropped if the models were built using more distant homologs, and we were able to quantitatively estimate this decline. However, the specificities of these techniques are still rather good even when the similarity is weak, which is a desirable characteristic For driving expensive Follow-up experiments. Our experiments suggest that though gene prediction will improve with every new protein that is discovered and through improvements in the current set of tools, we still have a long way to go before we can decipher the precise exonic structure of every gene in the human genome using purely computational methodology.