Annotated chemical patent corpus: a gold standard for text mining.

Annotated chemical patent corpus: a gold standard for text mining.
复制标题

DOI:
10.1371/journal.pone.0107477
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Muresan S
Muresan S
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Akhondi SA;Klenner AG;Tyrchan C;Manchala AK;Boppana K;Lowe D;Zimmermann M;Jagarlapudi SA;Sayle R;Kors JA;Muresan S

文献摘要

参考文献

被引文献

相似文献

探索专利申请所涵盖的化学和生物领域在早期药物化学活动中至关重要。专利分析可以提供对化合物现有技术的了解、新颖性检查、生物分析的验证,以及确定化学探索的新起点。通过专家馆长的人工提取从专利中提取化学和生物实体可能需要大量的时间和资源。文本挖掘方法可以帮助简化这一过程。为了验证这些方法的性能,人工标注的专利语料库是必不可少的。在本研究中,我们产生了一个大型的金标准化学专利语料库。我们制定了注释指南,并从世界知识产权组织、美国专利商标局和欧洲专利局挑选了200项完整专利。这些专利被自动预先注解,并提供给四个独立的注释者小组,每个小组由两到十个注释者组成。注释者在不同的亚类、疾病、目标和作用模式中标记了化学物质。还对由于光学字符识别错误而导致的拼写错误和伪换行符进行了注释。至少有三个注释员小组对47项专利的子集进行了注解,由此得出了统一的注解和注释者间协议分数。一组人对全套进行了注释。专利语料库包括全套专利的400,125条注释和协调集的36,537条注释。所有专利和注解实体均可在www.biosemantics.org上公开获得。
Exploring the chemical and biological space covered by patent applications is crucial in early-stage medicinal chemistry activities. Patent analysis can provide understanding of compound prior art, novelty checking, validation of biological assays, and identification of new starting points for chemical exploration. Extracting chemical and biological entities from patents through manual extraction by expert curators can take substantial amount of time and resources. Text mining methods can help to ease this process. To validate the performance of such methods, a manually annotated patent corpus is essential. In this study we have produced a large gold standard chemical patent corpus. We developed annotation guidelines and selected 200 full patents from the World Intellectual Property Organization, United States Patent and Trademark Office, and European Patent Office. The patents were pre-annotated automatically and made available to four independent annotator groups each consisting of two to ten annotators. The annotators marked chemicals in different subclasses, diseases, targets, and modes of action. Spelling mistakes and spurious line break due to optical character recognition errors were also annotated. A subset of 47 patents was annotated by at least three annotator groups, from which harmonized annotations and inter-annotator agreement scores were derived. One group annotated the full set. The patent corpus includes 400,125 annotations for the full set and 36,537 annotations for the harmonized set. All patents and annotated entities are publicly available at www.biosemantics.org.
DOI: 10.1021/ci200463r
发表时间: 2012-01-01
影响因子: 5.6
作者:
Sayle, Roger;Xie, Paul Hongxing;Muresan, Sorel
通讯作者: Muresan, Sorel
DOI: 10.1186/1758-2946-4-35
发表时间: 2012-12-13
影响因子: 8.6
作者:
Akhondi SA;Kors JA;Muresan S
通讯作者: Muresan S
DOI: 10.1186/1758-2946-3-14
发表时间: 2011-05-13
影响因子: 8.6
作者:
Southan C;Boppana K;Jagarlapudi SA;Muresan S
通讯作者: Muresan S
DOI: 10.1186/1758-2946-3-40
发表时间: 2011-10-14
影响因子: 8.6
作者:
Jessop DM;Adams SE;Murray-Rust P
通讯作者: Murray-Rust P
DOI: 10.1186/1758-2946-5-7
发表时间: 2013-01-24
影响因子: 8.6
作者:
Heller S;McNaught A;Stein S;Tchekhovskoi D;Pletnev I
通讯作者: Pletnev I