Mining protein function from text using term-based support vector machines.

Mining protein function from text using term-based support vector machines.
复制标题

DOI:
10.1186/1471-2105-6-s1-s22
复制
发表时间:
2005
期刊:
影响因子:
3
通讯作者:
Stapley BJ
Stapley BJ
中科院分区:
生物学4区
文献类型:
--
作者:
Rice SB;Nenadic G;Stapley BJ

文献摘要

被引文献

相似文献

文本挖掘激发了人们对生物学领域的巨大兴趣。BioCreAtIvE练习的目标是评估当前文本挖掘系统的性能。我们参加了任务2,该任务涉及将基因本体术语分配给人类蛋白质并从全文文档中选择相关证据。我们将其作为文档分类任务的修改形式来处理。我们使用有监督的机器学习方法(基于支持向量机)来分配蛋白质功能并选择支持分配的传代。作为分类特征,我们使用从文档中自动提取的蛋白质的共发生术语。管理员评估的结果是适度的,并且对于不同的问题变化很大:在许多情况下,我们对GO术语的蛋白质分配相对较好,但选择的支持文本通常是不相关的(精度从3%到50%不等)。当获得大量相关文件时,该方法似乎效果最好,而对于单个文件和/或简短段落则效果不佳。初步结果表明,我们的方法也可以从文本中挖掘注释,即使没有将蛋白质与GO术语相关的明确声明。从文本中挖掘蛋白质功能预测的机器学习方法只有在有足够的训练数据可用,并且有大量的支持数据用于预测的情况下才能产生良好的性能。最有希望的结果是结合文献检索和GO术语分配,这需要整合在BioCreAtIvE任务1和任务2中开发的方法。
Text mining has spurred huge interest in the domain of biology. The goal of the BioCreAtIvE exercise was to evaluate the performance of current text mining systems. We participated in Task 2, which addressed assigning Gene Ontology terms to human proteins and selecting relevant evidence from full-text documents. We approached it as a modified form of the document classification task. We used a supervised machine-learning approach (based on support vector machines) to assign protein function and select passages that support the assignments. As classification features, we used a protein's co-occurring terms that were automatically extracted from documents. The results evaluated by curators were modest, and quite variable for different problems: in many cases we have relatively good assignment of GO terms to proteins, but the selected supporting text was typically non-relevant (precision spanning from 3% to 50%). The method appears to work best when a substantial set of relevant documents is obtained, while it works poorly on single documents and/or short passages. The initial results suggest that our approach can also mine annotations from text even when an explicit statement relating a protein to a GO term is absent. A machine learning approach to mining protein function predictions from text can yield good performance only if sufficient training data is available, and significant amount of supporting data is used for prediction. The most promising results are for combined document retrieval and GO term assignment, which calls for the integration of methods developed in BioCreAtIvE Task 1 and Task 2.