A robust data-driven approach for gene ontology annotation.

A robust data-driven approach for gene ontology annotation.
复制标题

DOI:
10.1093/database/bau113
复制
发表时间:
2014
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Yu H
Yu H
中科院分区:
其他
文献类型:
--
作者:
Li Y;Yu H

文献摘要

参考文献

相似文献

基因本体(GO)和GO注释是生物信息管理和知识发现的重要资源,但手动注释的速度成为数据库管理的主要瓶颈。 BioCreative IV GO注释任务旨在评估根据生物医学文献中的叙述句子自动为基因分配GO术语的系统的性能。本文介绍了我们在这项任务中的工作以及竞赛后的实验结果。对于证据句子提取子任务,我们构建了一个二元分类器来使用参考距离估计器(RDE)来识别证据句子,这是一种最近提出的半监督学习方法,可以从大约 1000 万个未标记句子中学习新特征,在精确匹配中实现 19.3% 的 F1,在宽松匹配中实现 32.5% 的 F1。在提交后的实验中,我们通过在 RDE 学习中结合二元组特征,获得了 22.1% 和 35.7% 的 F1 性能。在开发和测试集中,基于 RDE 的方法在 F1 和 AUC 性能上相对于经典的监督学习方法(例如支持向量机和逻辑回归。对于 GO 术语预测子任务,我们开发了一种基于信息检索的方法,使用结合余弦相似度和文档中 GO 术语频率的排序函数以及基于高级 GO 类别的过滤方法来检索与每个证据句子最相关的 GO 术语。我们提交的运行的最佳性能是 7.8% F1 和 22.2% 层次结构 F1。我们发现频率信息和层次过滤的结合大大提高了性能。在提交后评估中,我们使用更简单的设置获得了 10.6% 的 F1。总体而言,实验分析表明我们的方法在这两项任务中都很稳健。
Gene ontology (GO) and GO annotation are important resources for biological information management and knowledge discovery, but the speed of manual annotation became a major bottleneck of database curation. BioCreative IV GO annotation task aims to evaluate the performance of system that automatically assigns GO terms to genes based on the narrative sentences in biomedical literature. This article presents our work in this task as well as the experimental results after the competition. For the evidence sentence extraction subtask, we built a binary classifier to identify evidence sentences using reference distance estimator (RDE), a recently proposed semi-supervised learning method that learns new features from around 10 million unlabeled sentences, achieving an F1 of 19.3% in exact match and 32.5% in relaxed match. In the post-submission experiment, we obtained 22.1% and 35.7% F1 performance by incorporating bigram features in RDE learning. In both development and test sets, RDE-based method achieved over 20% relative improvement on F1 and AUC performance against classical supervised learning methods, e.g. support vector machine and logistic regression. For the GO term prediction subtask, we developed an information retrieval-based method to retrieve the GO term most relevant to each evidence sentence using a ranking function that combined cosine similarity and the frequency of GO terms in documents, and a filtering method based on high-level GO classes. The best performance of our submitted runs was 7.8% F1 and 22.2% hierarchy F1. We found that the incorporation of frequency information and hierarchy filtering substantially improved the performance. In the post-submission evaluation, we obtained a 10.6% F1 using a simpler setting. Overall, the experimental analysis showed our approaches were robust in both the two tasks.
结合丰富的背景知识进行基因命名实体分类和识别
DOI: 10.1186/1471-2105-10-223
发表时间: 2009-07-17
期刊: BMC bioinformatics
影响因子: 3
作者:
Li Y;Lin H;Yang Z
通讯作者: Yang Z
DOI: 10.1109/tcbb.2010.99
发表时间: 2011-03-01
影响因子: 4.5
作者:
Li, Yanpeng;Hu, Xiaohua;Yang, Zhihao
通讯作者: Yang, Zhihao
DOI: 10.1186/1471-2105-6-s1-s2
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Yeh A;Morgan A;Colosimo M;Hirschman L
通讯作者: Hirschman L
BioCreative II 蛋白质-蛋白质相互作用注释提取任务概述。
DOI: 10.1186/gb-2008-9-s2-s4
发表时间: 2008
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Krallinger, Martin;Leitner, Florian;Rodriguez-Penagos, Carlos;Valencia, Alfonso
通讯作者: Valencia, Alfonso
DOI: 10.1186/gb-2008-9-s2-s2
发表时间: 2008
期刊: GENOME BIOLOGY
影响因子: 12.3
作者:
Smith, Larry;Tanabe, Lorraine K.;Johnson Nee Ando, Rie;Kuo, Cheng-Ju;Chung, I-Fang;Hsu, Chun-Nan;Lin, Yu-Shi;Klinger, Roman;Friedrich, Christoph M.;Ganchev, Kuzman;Torii, Manabu;Liu, Hongfang;Haddow, Barry;Struble, Craig A.;Povinelli, Richard J.;Vlachos, Andreas;Baumgartner, William A., Jr.;Hunter, Lawrence;Carpenter, Bob;Tsai, Richard Tzong-Han;Dai, Hong-Jie;Liu, Feng;Chen, Yifei;Sun, Chengjie;Katrenko, Sophia;Adriaans, Pieter;Blaschke, Christian;Torres, Rafael;Neves, Mariana;Nakov, Preslav;Divoli, Anna;Mana-Lopez, Manuel;Mata, Jacinto;Wilbur, W. John
通讯作者: Wilbur, W. John