The EU-ADR corpus: Annotated drugs, diseases, targets, and their relationships

The EU-ADR corpus: Annotated drugs, diseases, targets, and their relationships
复制标题

DOI:
10.1016/j.jbi.2012.04.004
复制
发表时间:
2012-10-01
影响因子:
4.5
通讯作者:
Furlong, Laura I.
Furlong, Laura I.
中科院分区:
医学3区
文献类型:
--
作者:
van Mulligen, Erik M.;Fourrier-Reglat, Annie;Furlong, Laura I.

文献摘要

被引文献

相似文献

标注了特定实体和关系的语料库对于训练和评估文本挖掘系统是必不可少的,这些系统是为了从大型语料库中提取特定的结构化信息而开发的。在本文中,我们描述了一种方法,其中命名实体识别系统产生的第一个注释和注释者使用基于Web的接口修改此注释。所获得的一致性数据表明,注释者之间的一致性要比与系统提供的注释的一致性好得多。该语料库已被注释为药物,疾病,基因及其相互关系。对于每一种药物-病症、药物-靶点和靶点-病症关系,三位专家注释了一组100篇摘要。这些注释关系将用于训练和评估文本挖掘软件,以捕获文本中的这些关系。(C)2012 Elsevier Inc. All rights reserved.
Corpora with specific entities and relationships annotated are essential to train and evaluate text-mining systems that are developed to extract specific structured information from a large corpus. In this paper we describe an approach where a named-entity recognition system produces a first annotation and annotators revise this annotation using a web-based interface. The agreement figures achieved show that the inter-annotator agreement is much better than the agreement with the system provided annotations. The corpus has been annotated for drugs, disorders, genes and their inter-relationships. For each of the drug-disorder, drug-target, and target-disorder relations three experts have annotated a set of 100 abstracts. These annotated relationships will be used to train and evaluate text-mining software to capture these relationships in texts. (C) 2012 Elsevier Inc. All rights reserved.