Rutabaga by any other name: extracting biological names

Rutabaga by any other name: extracting biological names
复制标题

DOI:
10.1016/s1532-0464(03)00014-5
复制
发表时间:
2002-08-01
影响因子:
4.5
通讯作者:
Yeh, AS
Yeh, AS
中科院分区:
医学3区
文献类型:
--
作者:
Hirschman, L;Morgan, AA;Yeh, AS

文献摘要

被引文献

相似文献

随着生物学研究步伐的加快,生物学家正变得越来越依赖计算机来管理信息爆炸。生物学家通过依赖精确的生物学术语来交流他们的研究结果;这些术语然后提供进入文献和越来越多的生物学数据库的索引。本文研究了通过提取实体名称和它们之间的关系来获取生物资源的新兴技术。信息抽取一直是自然语言处理领域的一个活跃研究领域,将信息抽取应用于新闻故事也取得了很好的结果,例如,在识别人名、组织名和地名方面,信息抽取的准确率和召回率达到了93%-95%。但这些结果似乎不会直接转移到生物名称上,生物名称的结果保持在75%-80%的范围内。可能涉及多种因素,包括缺乏用于严格衡量进展的共享训练和测试集、缺乏专门针对生物任务的带注释的训练数据、术语普遍含糊、频繁引入新术语以及为新闻和真实生物问题定义的评价任务之间的不匹配。我们提供了一个简单的词汇匹配练习的证据,该练习说明了在识别生物名称时遇到的一些具体问题。最后,我们概述了一项研究议程,旨在将命名实体标记的性能提高到可用于执行具有生物学重要性的任务的水平。(C)2003年埃尔塞维尔科学公司(美国)。版权所有。
As the pace of biological research accelerates, biologists are becoming increasingly reliant on computers to manage the information explosion. Biologists communicate their research findings by relying on precise biological terms; these terms then provide indices into the literature and across the growing number of biological databases. This article examines emerging techniques to access biological resources through extraction of entity names and relations among them. Information extraction has been an active area of research in natural language processing and there are promising results for information extraction applied to news stories, e.g., balanced precision and recall in the 93-95% range for identifying person, organization and location names. But these results do not seem to transfer directly to biological names, where results remain in the 75-80% range. Multiple factors may be involved, including absence of shared training and test sets for rigorous measures of progress, lack of annotated training data specific to biological tasks, pervasive ambiguity of terms, frequent introduction of new terms, and a mismatch between evaluation tasks as defined for news and real biological problems. We present evidence from a simple lexical matching exercise that illustrates some specific problems encountered when identifying biological names. We conclude by outlining a research agenda to raise performance of named entity tagging to a level where it can be used to perform tasks of biological importance. (C) 2003 Elsevier Science (USA). All rights reserved.