Wide coverage biomedical event extraction using multiple partially overlapping corpora.

Wide coverage biomedical event extraction using multiple partially overlapping corpora.
复制标题

DOI:
10.1186/1471-2105-14-175
复制
发表时间:
2013-06-03
期刊:
影响因子:
3
通讯作者:
Ananiadou S
Ananiadou S
中科院分区:
生物学4区
文献类型:
--
作者:
Miwa M;Pyysalo S;Ohta T;Ananiadou S

文献摘要

参考文献

被引文献

相似文献

生物医学事件是理解生理过程和疾病的关键,并且需要广泛覆盖的提取来全面自动分析文献中描述生物医学系统的陈述。反过来,提取方法的训练和评估需要手动注释的语料库。然而,由于人工标注是耗时和昂贵的,任何单一的事件标注语料库只能覆盖有限数量的语义类型。虽然结合使用几个这样的语料库可能会允许提取系统,以实现广泛的语义覆盖,很少有研究学习多个语料库与部分重叠的语义注释范围。我们提出了一种从多个语料库中学习部分语义标注重叠的方法,并实现了这种方法,以改善我们现有的事件提取系统,EventMine。使用七个事件注释语料库进行评估,总共包括65种事件类型,表明从重叠语料库中学习可以产生一个单一的,语料库独立的,广泛覆盖的提取系统,其性能优于在单个语料库上训练的系统,并超过了先前在BioNLP共享任务2011中两个已建立的事件提取任务上报告的结果。所提出的方法允许从多个语料库中训练一个覆盖范围广,最先进的事件提取系统,部分语义注释重叠。由此产生的单个模型在实践中通过消除选择兼容语料库或语义类型的子集或合并在不同个体语料库上训练的多个模型的结果的需要,使广泛覆盖的提取变得简单。多语料库学习还允许注释工作专注于覆盖额外的语义类型,而不是在任何单个注释工作中进行详尽的覆盖,或者扩展现有语料库中注释的语义类型的覆盖范围。
Biomedical events are key to understanding physiological processes and disease, and wide coverage extraction is required for comprehensive automatic analysis of statements describing biomedical systems in the literature. In turn, the training and evaluation of extraction methods requires manually annotated corpora. However, as manual annotation is time-consuming and expensive, any single event-annotated corpus can only cover a limited number of semantic types. Although combined use of several such corpora could potentially allow an extraction system to achieve broad semantic coverage, there has been little research into learning from multiple corpora with partially overlapping semantic annotation scopes. We propose a method for learning from multiple corpora with partial semantic annotation overlap, and implement this method to improve our existing event extraction system, EventMine. An evaluation using seven event annotated corpora, including 65 event types in total, shows that learning from overlapping corpora can produce a single, corpus-independent, wide coverage extraction system that outperforms systems trained on single corpora and exceeds previously reported results on two established event extraction tasks from the BioNLP Shared Task 2011. The proposed method allows the training of a wide-coverage, state-of-the-art event extraction system from multiple corpora with partial semantic annotation overlap. The resulting single model makes broad-coverage extraction straightforward in practice by removing the need to either select a subset of compatible corpora or semantic types, or to merge results from several models trained on different individual corpora. Multi-corpus learning also allows annotation efforts to focus on covering additional semantic types, rather than aiming for exhaustive coverage in any single annotation effort, or extending the coverage of semantic types annotated in existing corpora.
DOI: 10.1186/1471-2105-13-s11-s9
发表时间: 2012-06-26
期刊: BMC bioinformatics
影响因子: 3
作者:
McClosky D;Riedel S;Surdeanu M;McCallum A;Manning CD
通讯作者: Manning CD
DOI: 10.1613/jair.1872
发表时间: 2006-01-01
影响因子: 5
作者:
Daumé, H;Marcu, D
通讯作者: Marcu, D
DOI: 10.1186/1471-2105-9-10
发表时间: 2008-01-08
期刊: BMC bioinformatics
影响因子: 3
作者:
Kim JD;Ohta T;Tsujii J
通讯作者: Tsujii J
DOI: 10.1038/msb.2010.108
发表时间: 2010-12-21
影响因子: 9.9
作者:
Caron, Etienne;Ghosh, Samik;Matsuoka, Yukiko;Ashton-Beaucage, Dariel;Therrien, Marc;Lemieux, Sebastien;Perreault, Claude;Roux, Philippe P.;Kitano, Hiroaki
通讯作者: Kitano, Hiroaki
DOI: 10.1186/1471-2105-13-s11-s2
发表时间: 2012-06-26
期刊: BMC bioinformatics
影响因子: 3
作者:
Pyysalo S;Ohta T;Rak R;Sullivan D;Mao C;Wang C;Sobral B;Tsujii J;Ananiadou S
通讯作者: Ananiadou S