The CHEMDNER corpus of chemicals and drugs and its annotation principles.

The CHEMDNER corpus of chemicals and drugs and its annotation principles.
复制标题

DOI:
10.1186/1758-2946-7-s1-s2
复制
发表时间:
2015
影响因子:
8.6
通讯作者:
Valencia A
Valencia A
中科院分区:
化学2区
文献类型:
--
作者:
Krallinger M;Rabal O;Leitner F;Vazquez M;Salgado D;Lu Z;Leaman R;Lu Y;Ji D;Lowe DM;Sayle RA;Batista-Navarro RT;Rak R;Huber T;Rocktäschel T;Matos S;Campos D;Tang B;Xu H;Munkhdalai T;Ryu KH;Ramanan SV;Nathan S;Žitnik S;Bajec M;Weber L;Irmer M;Akhondi SA;Kors JA;Xu S;An X;Sikdar UK;Ekbal A;Yoshioka M;Dieb TM;Choi M;Verspoor K;Khabsa M;Giles CL;Liu H;Ravikumar KE;Lamurias A;Couto FM;Dai HJ;Tsai RT;Ata C;Can T;Usié A;Alves R;Segura-Bedmar I;Martínez P;Oyarzabal J;Valencia A

文献摘要

被引文献

相似文献

从文本中自动提取化学信息需要识别化学实体提及作为其关键步骤之一。在开发有监督的命名实体识别 (NER) 系统时,需要有一个大型的、手动注释的文本语料库。此外,大型语料库允许对检测文档中化学物质的不同方法进行稳健的评估和比较。我们展示了 CHEMDNER 语料库,它是 10,000 条 PubMed 摘要的集合,其中总共包含 84,355 个化学实体提及项,由化学专家文献管理员手动标记,遵循专门为此任务定义的注释指南。 CHEMDNER 语料库的摘要被选择为代表所有主要化学学科。每个化学实体提及都是根据其结构相关化学实体提及 (SACEM) 类别手动标记的:缩写、家族、公式、标识符、多重、系统和琐碎。使用注释者之间的一致性研究来测量文本中标记化学品的难度和一致性,获得了 91 的一致性百分比。对于 CHEMDNER 语料库的子集(3,000 个摘要的测试集),我们不仅提供黄金标准手动注释,而且还提供由参与 BioCreative IV CHEMDNER 化学提及识别任务的 26 个团队自动检测到的提及。此外,我们还发布了 CHEMDNER 银标准语料库,其中包含从 17,000 篇随机选择的 PubMed 摘要中自动提取的提及内容。 BioC 格式的 CHEMDNER 语料库版本也已生成。我们提出了一个关于实体注释所需的最低信息的标准,用于构建化学和药物实体的领域特定语料库。 CHEMDNER 语料库和注释指南可在以下网址获取:http://www.biocreative.org/resources/biocreative-iv/chemdner-corpus/
The automatic extraction of chemical information from text requires the recognition of chemical entity mentions as one of its key steps. When developing supervised named entity recognition (NER) systems, the availability of a large, manually annotated text corpus is desirable. Furthermore, large corpora permit the robust evaluation and comparison of different approaches that detect chemicals in documents. We present the CHEMDNER corpus, a collection of 10,000 PubMed abstracts that contain a total of 84,355 chemical entity mentions labeled manually by expert chemistry literature curators, following annotation guidelines specifically defined for this task. The abstracts of the CHEMDNER corpus were selected to be representative for all major chemical disciplines. Each of the chemical entity mentions was manually labeled according to its structure-associated chemical entity mention (SACEM) class: abbreviation, family, formula, identifier, multiple, systematic and trivial. The difficulty and consistency of tagging chemicals in text was measured using an agreement study between annotators, obtaining a percentage agreement of 91. For a subset of the CHEMDNER corpus (the test set of 3,000 abstracts) we provide not only the Gold Standard manual annotations, but also mentions automatically detected by the 26 teams that participated in the BioCreative IV CHEMDNER chemical mention recognition task. In addition, we release the CHEMDNER silver standard corpus of automatically extracted mentions from 17,000 randomly selected PubMed abstracts. A version of the CHEMDNER corpus in the BioC format has been generated as well. We propose a standard for required minimum information about entity annotations for the construction of domain specific corpora on chemical and drug entities. The CHEMDNER corpus and annotation guidelines are available at: http://www.biocreative.org/resources/biocreative-iv/chemdner-corpus/