Improving the Caenorhabditis elegans genome annotation using machine learning.

Improving the Caenorhabditis elegans genome annotation using machine learning.
复制标题

DOI:
10.1371/journal.pcbi.0030020
复制
发表时间:
2007-02-23
影响因子:
4.3
通讯作者:
Schölkopf B
Schölkopf B
中科院分区:
生物学2区
文献类型:
--
作者:
Rätsch G;Sonnenburg S;Srinivasan J;Witte H;Müller KR;Sommer RJ;Schölkopf B

文献摘要

参考文献

被引文献

相似文献

对于现代生物学来说,精确的基因组注释是至关重要的,因为它们允许基因区域的准确定义。我们采用最先进的机器学习方法来分析和提高线虫线虫的基因组注释的准确性。建议的机器学习系统进行训练,以识别未剪接的mRNA上的外显子和内含子,利用支持向量机和标签序列学习的最新进展。在87%(编码区和非翻译区)和95%(仅编码区)的所有基因在几个样本外的评估测试,我们的方法正确地确定了所有的外显子和内含子。值得注意的是,在C.线虫基因组注释与我们的预测一致,因此我们假设这些基因中有相当大的一部分没有被正确注释。对C. elegans的研究表明,对WS 120中未经证实的基因的剪接形式预测在约18%的考虑情况下是不准确的,而我们的预测仅在10%-13%的情况下偏离事实。我们实验分析了20个有争议的基因,我们的系统和注释不一致,证实了我们的预测的优越性。虽然我们的方法正确预测了75%的情况,但标准注释从未完全正确。我们的系统的准确性进一步证实了与其他两个最近提出的系统,可用于剪接形式预测:SNAP和ExonHunter的比较。我们认为C.使用现代机器学习技术可以大大增强线虫和其他生物体。真核基因含有内含子,内含子是从基因转录物中切除的插入序列,伴随着称为外显子的侧翼片段的连接。去除内含子的过程称为剪接。它涉及到迄今为止过于复杂的生化机制,无法全面准确地建模。然而,丰富的测序结果可以作为一个蓝图数据库,举例说明这个过程完成。使用这个数据库,我们采用判别式机器学习技术来预测未剪接的前mRNA的成熟mRNA。我们的方法利用支持向量机和标签序列学习的最新进展,最初是为自然语言处理开发的。该系统名为mSplicer,在线虫C的基因组上进行了训练和评估。一种研究得很好的模式生物。我们能够证明mSplicer在大多数情况下正确预测了拼接形式。令人惊讶的是,我们对目前未经证实的基因的预测与公开的基因组注释有很大的偏差。据推测,这些基因中有相当大的一部分没有被正确注释。回顾性评估和额外的测序结果显示了mSplicer预测的优越性。可以得出结论,使用现代机器学习可以大大增强线虫和其他基因组的注释。
For modern biology, precise genome annotations are of prime importance, as they allow the accurate definition of genic regions. We employ state-of-the-art machine learning methods to assay and improve the accuracy of the genome annotation of the nematode Caenorhabditis elegans. The proposed machine learning system is trained to recognize exons and introns on the unspliced mRNA, utilizing recent advances in support vector machines and label sequence learning. In 87% (coding and untranslated regions) and 95% (coding regions only) of all genes tested in several out-of-sample evaluations, our method correctly identified all exons and introns. Notably, only 37% and 50%, respectively, of the presently unconfirmed genes in the C. elegans genome annotation agree with our predictions, thus we hypothesize that a sizable fraction of those genes are not correctly annotated. A retrospective evaluation of the Wormbase WS120 annotation of C. elegans reveals that splice form predictions on unconfirmed genes in WS120 are inaccurate in about 18% of the considered cases, while our predictions deviate from the truth only in 10%–13%. We experimentally analyzed 20 controversial genes on which our system and the annotation disagree, confirming the superiority of our predictions. While our method correctly predicted 75% of those cases, the standard annotation was never completely correct. The accuracy of our system is further corroborated by a comparison with two other recently proposed systems that can be used for splice form prediction: SNAP and ExonHunter. We conclude that the genome annotation of C. elegans and other organisms can be greatly enhanced using modern machine learning technology. Eukaryotic genes contain introns, which are intervening sequences that are excised from a gene transcript with the concomitant ligation of flanking segments called exons. The process of removing introns is called splicing. It involves biochemical mechanisms that to date are too complex to be modeled comprehensively and accurately. However, abundant sequencing results can serve as a blueprint database exemplifying what this process accomplishes. Using this database, we employ discriminative machine learning techniques to predict the mature mRNA given the unspliced pre-mRNA. Our method utilizes support vector machines and recent advances in label sequence learning, originally developed for natural language processing. The system, called mSplicer, was trained and evaluated on the genome of the nematode C. elegans, a well-studied model organism. We were able to show that mSplicer correctly predicts the splice form in most cases. Surprisingly, our predictions on currently unconfirmed genes deviate considerably from the public genome annotation. It is hypothesized that a sizable fraction of those genes are not correctly annotated. A retrospective evaluation and additional sequencing results show the superiority of mSplicer's predictions. It is concluded that the annotation of nematode and other genomes can be greatly enhanced using modern machine learning.
DOI: 10.1073/pnas.97.1.262
发表时间: 2000-01-04
影响因子: 11.1
作者:
Brown, MPS;Grundy, WN;Haussler, D
通讯作者: Haussler, D
DOI: 10.1101/gr.10.4.529
发表时间: 2000-04-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Reese, MG;Kulp, D;Haussler, D
通讯作者: Haussler, D
DOI: 10.1093/nar/gkg359
发表时间: 2003-05-15
影响因子: 14.9
作者:
Lee, KZ;Eizinger, A;Sommer, RJ
通讯作者: Sommer, RJ
DOI: 10.1016/j.scico.2003.12.005
发表时间: 2004-06-01
影响因子: 1.3
作者:
Giegerich, R;Meyer, C;Steffen, P
通讯作者: Steffen, P
DOI: 10.1073/pnas.76.3.1333
发表时间: 1979-01-01
影响因子: 11.1
作者:
EMMONS, SW;KLASS, MR;HIRSH, D
通讯作者: HIRSH, D