Automating curation using a natural language processing pipeline.

Automating curation using a natural language processing pipeline.
复制标题

DOI:
10.1186/gb-2008-9-s2-s10
复制
发表时间:
2008
期刊:
影响因子:
12.3
通讯作者:
Wang X
Wang X
中科院分区:
生物学1区
文献类型:
--
作者:
Alex B;Grover C;Haddow B;Kabadjov M;Klein E;Matthews M;Tobin R;Wang X

文献摘要

参考文献

被引文献

相似文献

BioCreative II中的任务旨在近似于策划生物医学研究论文所涉及的一些艰苦工作。爱丁堡大学团队完成这些任务的方法是调整和扩展现有的自然语言处理(NLP)系统,该系统是我们作为商业策展助理的一部分开发的。虽然本文集中于使用NLP来协助策展,但该系统同样可以用于从文献中提取与生物学家直接相关的信息类型。我们的系统是最高的执行互动子任务,并在基因提及任务的竞争力表现是以最小的发展努力。对于基因归一化任务,可以快速应用于新领域的字符串匹配技术显示出接近平均水平的表现。正在开发的技术被证明是很容易适应BioCreative II的任务。虽然高性能可以在单个任务上获得,例如基因提及识别和标准化,以及文档分类,但必须组合许多组件的任务,例如相互作用蛋白质对的检测和标准化,对于NLP系统仍然具有挑战性。
The tasks in BioCreative II were designed to approximate some of the laborious work involved in curating biomedical research papers. The approach to these tasks taken by the University of Edinburgh team was to adapt and extend the existing natural language processing (NLP) system that we have developed as part of a commercial curation assistant. Although this paper concentrates on using NLP to assist with curation, the system can be equally employed to extract types of information from the literature that is immediately relevant to biologists in general. Our system was among the highest performing on the interaction subtasks, and competitive performance on the gene mention task was achieved with minimal development effort. For the gene normalization task, a string matching technique that can be quickly applied to new domains was shown to perform close to average. The technologies being developed were shown to be readily adapted to the BioCreative II tasks. Although high performance may be obtained on individual tasks such as gene mention recognition and normalization, and document classification, tasks in which a number of components must be combined, such as detection and normalization of interacting protein pairs, are still challenging for NLP systems.
DOI: 10.1016/j.jbi.2004.08.008
发表时间: 2004-12-01
影响因子: 4.5
作者:
Collier, N;Takeuchi, K
通讯作者: Takeuchi, K
评估生物学的文本挖掘系统:第二次生物综合社区挑战的概述。
DOI: 10.1186/gb-2008-9-s2-s1
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Krallinger M;Morgan A;Smith L;Leitner F;Tanabe L;Wilbur J;Hirschman L;Valencia A
通讯作者: Valencia A
DOI: 10.1371/journal.pbio.0030065
发表时间: 2005-02
期刊: PLoS biology
影响因子: 9.8
作者:
Rebholz-Schuhmann D;Kirsch H;Couto F
通讯作者: Couto F
DOI: 10.1093/nar/28.1.45
发表时间: 2000-01-01
影响因子: 14.9
作者:
Bairoch, A;Apweiler, R
通讯作者: Apweiler, R
DOI: 10.2307/2289924
发表时间: 1989-06-01
影响因子: 3.7
作者:
JARO, MA
通讯作者: JARO, MA