Recognizing chemicals in patents: a comparative analysis.

Recognizing chemicals in patents: a comparative analysis.
复制标题

DOI:
10.1186/s13321-016-0172-0
复制
发表时间:
2016
影响因子:
8.6
通讯作者:
Leser U
Leser U
中科院分区:
化学2区
文献类型:
--
作者:
Habibi M;Wiegandt DL;Schmedding F;Leser U

文献摘要

参考文献

被引文献

相似文献

最近,化学命名实体识别(NER)的方法已经获得了很大的兴趣,自动分析的需要,今天不断增长的生物医学文本的集合驱动。由于药物发现的高度经济重要性,专利的化学NER特别重要。然而,由于缺乏足够的注释语料库,专利净认率长期以来基本上被研究界所忽视。最近的一项国际竞赛专门针对这一任务,但只对黄金标准专利摘要而不是完整专利进行评估;此外,由于训练和测试数据的同质性相对较高,此类竞赛的结果通常难以外推到现实生活中。在这里,我们评估了两个国家的最先进的化学NER工具,tmChem和ChemSpot,在四个不同的注释专利语料库,其中两个包括全文。我们研究了工具的整体性能,在实例级别上比较了它们的结果,报告了高召回率和高精度的集合,并进行了跨语料库和语料库内的评估。我们的研究结果表明,完整的专利比专利摘要更难分析,并清楚地证实了使用相同的文本类型(专利与科学)和文本类型(摘要与全文)进行训练和测试是实现高质量文本挖掘结果的先决条件。
Recently, methods for Chemical Named Entity Recognition (NER) have gained substantial interest, driven by the need for automatically analyzing todays ever growing collections of biomedical text. Chemical NER for patents is particularly essential due to the high economic importance of pharmaceutical findings. However, NER on patents has essentially been neglected by the research community for long, mostly because of the lack of enough annotated corpora. A recent international competition specifically targeted this task, but evaluated tools only on gold standard patent abstracts instead of full patents; furthermore, results from such competitions are often difficult to extrapolate to real-life settings due to the relatively high homogeneity of training and test data. Here, we evaluate the two state-of-the-art chemical NER tools, tmChem and ChemSpot, on four different annotated patent corpora, two of which consist of full texts. We study the overall performance of the tools, compare their results at the instance level, report on high-recall and high-precision ensembles, and perform cross-corpus and intra-corpus evaluations. Our findings indicate that full patents are considerably harder to analyze than patent abstracts and clearly confirm the common wisdom that using the same text genre (patent vs. scientific) and text type (abstract vs. full text) for training and testing is a pre-requisite for achieving high quality text mining results.
DOI: 10.1186/1471-2105-13-161
发表时间: 2012-07-09
期刊: BMC bioinformatics
影响因子: 3
作者:
Bada M;Eckert M;Evans D;Garcia K;Shipley K;Sitnikov D;Baumgartner WA Jr;Cohen KB;Verspoor K;Blake JA;Hunter LE
通讯作者: Hunter LE
DOI: 10.1186/1758-2946-7-s1-s2
发表时间: 2015
影响因子: 8.6
作者:
Krallinger M;Rabal O;Leitner F;Vazquez M;Salgado D;Lu Z;Leaman R;Lu Y;Ji D;Lowe DM;Sayle RA;Batista-Navarro RT;Rak R;Huber T;Rocktäschel T;Matos S;Campos D;Tang B;Xu H;Munkhdalai T;Ryu KH;Ramanan SV;Nathan S;Žitnik S;Bajec M;Weber L;Irmer M;Akhondi SA;Kors JA;Xu S;An X;Sikdar UK;Ekbal A;Yoshioka M;Dieb TM;Choi M;Verspoor K;Khabsa M;Giles CL;Liu H;Ravikumar KE;Lamurias A;Couto FM;Dai HJ;Tsai RT;Ata C;Can T;Usié A;Alves R;Segura-Bedmar I;Martínez P;Oyarzabal J;Valencia A
通讯作者: Valencia A
DOI: 10.1186/1758-2946-3-17
发表时间: 2011-05-16
影响因子: 8.6
作者:
Hawizy L;Jessop DM;Adams N;Murray-Rust P
通讯作者: Murray-Rust P
DOI: 10.1186/1758-2946-7-s1-s3
发表时间: 2015
影响因子: 8.6
作者:
Leaman R;Wei CH;Lu Z
通讯作者: Lu Z
DOI: 10.1093/bioinformatics/bts183
发表时间: 2012-06-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rocktaschel, Tim;Weidlich, Michael;Leser, Ulf
通讯作者: Leser, Ulf