BioCause: Annotating and analysing causality in the biomedical domain.

BioCause: Annotating and analysing causality in the biomedical domain.
复制标题

DOI:
10.1186/1471-2105-14-2
复制
发表时间:
2013-01-16
期刊:
影响因子:
3
通讯作者:
Ananiadou S
Ananiadou S
中科院分区:
生物学4区
文献类型:
--
作者:
Mihăilă C;Ohta T;Pyysalo S;Ananiadou S

文献摘要

参考文献

被引文献

相似文献

带有事件级信息注释的生物医学语料库是特定领域信息提取(IE)系统的重要资源。然而,生物事件注释本身并不能满足生物学家的所有需求。与关系和事件提取的工作不同,大多数工作集中在特定事件和命名实体上,我们的目标是建立一个全面的资源,涵盖话语中存在的因果关联的所有陈述。因果关系是生物医学知识(如诊断、病理学或系统生物学)的核心,因此,自动因果关系识别可以通过提示可能的因果关系和帮助管理途径模型,大大减少人类的工作量。因此,带有这种关系的生物医学文本语料库对于开发和评估生物医学文本挖掘至关重要。我们定义了一种用因果关系丰富生物医学领域语料库的标注方案。该模式随后被用于注释851个因果关系,形成BioCause,这是一个属于传染病子领域的19篇开放获取的生物医学期刊全文文章的集合。在之前的共享任务上下文中,这些文档已经用命名实体和事件信息进行了预先注释。我们报告了使用精确匹配约束的触发器的注释器间协议率超过60%,参数的注释器间协议率超过80%。使用宽松的匹配设置,这些会显著增加。此外,我们还从多个角度分析和描述了《生物成因》中的因果关系。这些信息可以用来训练自动因果关系检测系统。用关于因果话语关系的信息增强命名实体和事件注释有助于开发更复杂的IE系统。这将进一步影响多种任务的发展,例如使文本推理能够检测蕴涵,发现新的事实并为实验工作提供新的假设。
Biomedical corpora annotated with event-level information represent an important resource for domain-specific information extraction (IE) systems. However, bio-event annotation alone cannot cater for all the needs of biologists. Unlike work on relation and event extraction, most of which focusses on specific events and named entities, we aim to build a comprehensive resource, covering all statements of causal association present in discourse. Causality lies at the heart of biomedical knowledge, such as diagnosis, pathology or systems biology, and, thus, automatic causality recognition can greatly reduce the human workload by suggesting possible causal connections and aiding in the curation of pathway models. A biomedical text corpus annotated with such relations is, hence, crucial for developing and evaluating biomedical text mining. We have defined an annotation scheme for enriching biomedical domain corpora with causality relations. This schema has subsequently been used to annotate 851 causal relations to form BioCause, a collection of 19 open-access full-text biomedical journal articles belonging to the subdomain of infectious diseases. These documents have been pre-annotated with named entity and event information in the context of previous shared tasks. We report an inter-annotator agreement rate of over 60% for triggers and of over 80% for arguments using an exact match constraint. These increase significantly using a relaxed match setting. Moreover, we analyse and describe the causality relations in BioCause from various points of view. This information can then be leveraged for the training of automatic causality detection systems. Augmenting named entity and event annotations with information about causal discourse relations could benefit the development of more sophisticated IE systems. These will further influence the development of multiple tasks, such as enabling textual inference to detect entailments, discovering new facts and providing new hypotheses for experimental work.
DOI: 10.1016/j.jbi.2011.07.001
发表时间: 2011-12
影响因子: 4.5
作者:
Kleinberg, Samantha;Hripcsak, George
通讯作者: Hripcsak, George
DOI: 10.1371/journal.pcbi.0040020
发表时间: 2008-01
影响因子: 4.3
作者:
Cohen KB;Hunter L
通讯作者: Hunter L
DOI: 10.1093/bioinformatics/btg015
发表时间: 2003-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hucka, M;Finney, A;Wang, J
通讯作者: Wang, J
DOI: 10.1186/1471-2105-9-s11-s10
发表时间: 2008-11-19
期刊: BMC bioinformatics
影响因子: 3
作者:
Kilicoglu H;Bergler S
通讯作者: Bergler S
DOI: 10.1186/1471-2105-9-10
发表时间: 2008-01-08
期刊: BMC bioinformatics
影响因子: 3
作者:
Kim JD;Ohta T;Tsujii J
通讯作者: Tsujii J