A realistic assessment of methods for extracting gene/protein interactions from free text.

A realistic assessment of methods for extracting gene/protein interactions from free text.
复制标题

DOI:
10.1186/1471-2105-10-233
复制
发表时间:
2009-07-28
期刊:
影响因子:
3
通讯作者:
Shepherd AJ
Shepherd AJ
中科院分区:
生物学4区
文献类型:
--
作者:
Kabiljo R;Clegg AB;Shepherd AJ

文献摘要

参考文献

被引文献

相似文献

从文献中自动提取基因和/或蛋白质相互作用是生物医学文本挖掘研究的重要目标之一。在这篇文章中,我们提出了一个与潜在的非专家用户相关的基因/蛋白质交互挖掘的现实评估。因此,我们特别避免了安装复杂或需要重新实现的方法,并将我们选择的提取方法与最先进的生物医学命名实体标记器结合在一起。我们的结果表明:不同评估语料库的性能差异极大;标记(而不是黄金标准)基因和蛋白质名称的使用对性能有显著影响,F分数下降超过20个百分点是司空见惯的;当与命名实体标记相结合时,基于关键字的简单基准算法的性能优于两种最广泛用于提取基因/蛋白质相互作用的工具。在可用性、易用性和性能方面,对从自由文本中自动提取基因和/或蛋白质相互作用感兴趣的潜在非专业用户群体,目前的工具和系统服务不佳。公开发布易于安装和使用的提取工具,并达到最先进的性能水平,应该被生物医学文本挖掘社区视为高度优先的事项。
The automated extraction of gene and/or protein interactions from the literature is one of the most important targets of biomedical text mining research. In this paper we present a realistic evaluation of gene/protein interaction mining relevant to potential non-specialist users. Hence we have specifically avoided methods that are complex to install or require reimplementation, and we coupled our chosen extraction methods with a state-of-the-art biomedical named entity tagger. Our results show: that performance across different evaluation corpora is extremely variable; that the use of tagged (as opposed to gold standard) gene and protein names has a significant impact on performance, with a drop in F-score of over 20 percentage points being commonplace; and that a simple keyword-based benchmark algorithm when coupled with a named entity tagger outperforms two of the tools most widely used to extract gene/protein interactions. In terms of availability, ease of use and performance, the potential non-specialist user community interested in automatically extracting gene and/or protein interactions from free text is poorly served by current tools and systems. The public release of extraction tools that are easy to install and use, and that achieve state-of-art levels of performance should be treated as a high priority by the biomedical text mining community.
DOI: 10.1093/bioinformatics/btm557
发表时间: 2008-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rebholz-Schuhmann, Dietrich;Arregui, Miguel;Jimeno, Antonio
通讯作者: Jimeno, Antonio
DOI: 10.1186/1471-2105-9-10
发表时间: 2008-01-08
期刊: BMC bioinformatics
影响因子: 3
作者:
Kim JD;Ohta T;Tsujii J
通讯作者: Tsujii J
评估生物学的文本挖掘系统:第二次生物综合社区挑战的概述。
DOI: 10.1186/gb-2008-9-s2-s1
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Krallinger M;Morgan A;Smith L;Leitner F;Tanabe L;Wilbur J;Hirschman L;Valencia A
通讯作者: Valencia A
DOI: 10.1016/j.artmed.2004.07.016
发表时间: 2005-02-01
影响因子: 7.5
作者:
Bunescu, R;Ge, RF;Wong, YW
通讯作者: Wong, YW
DOI: 10.1093/bioinformatics/bti475
发表时间: 2005-07-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Settles, B
通讯作者: Settles, B