Optimising chemical named entity recognition with pre-processing analytics, knowledge-rich features and heuristics.

Optimising chemical named entity recognition with pre-processing analytics, knowledge-rich features and heuristics.
复制标题

DOI:
10.1186/1758-2946-7-s1-s6
复制
发表时间:
2015
影响因子:
8.6
通讯作者:
Ananiadou S
Ananiadou S
中科院分区:
化学2区
文献类型:
--
作者:
Batista-Navarro R;Rak R;Ananiadou S

文献摘要

被引文献

相似文献

化学命名实体识别(一项具有挑战性的自然语言处理任务)的强大方法的开发之前因缺乏公开可用的大规模金标准语料库而受到阻碍。最近公开发布的大型化学实体注释语料库作为第四届生物创新挑战评估(BioCreative IV)研讨会CHEMDNER轨道的资源,大大缓解了这个问题,并使我们能够开发一个有条件的随机字段为基础的化学实体识别器。为了优化其性能,我们在解决方案的各个方面引入了定制。其中包括选择专门的预处理分析,在统计模型的训练和应用中纳入化学知识丰富的功能,以及添加后处理规则。我们的评估表明,当我们的定制集成到化学实体识别器中时,可以获得最佳性能。当将其性能与最先进方法的性能进行比较时,在可比的实验设置下,我们的解决方案具有竞争优势。我们还表明,我们的识别器使用的模型训练的CHEMDNER语料库是适合识别的名称在广泛的语料库,始终优于两个流行的化学NER工具。这项工作的贡献是双重的。首先,我们提出了一个化学实体识别方法的细节,该方法已被证明具有竞争力的性能,如果不是上级,作为国家的最先进的方法的水平。其次,已开发的解决方案套件已公开可作为一个可配置的工作流程,在互操作的文本挖掘工作台Argo。这使得感兴趣的用户可以在其他化学文本挖掘任务的背景下方便地应用和评估我们的解决方案。
The development of robust methods for chemical named entity recognition, a challenging natural language processing task, was previously hindered by the lack of publicly available, large-scale, gold standard corpora. The recent public release of a large chemical entity-annotated corpus as a resource for the CHEMDNER track of the Fourth BioCreative Challenge Evaluation (BioCreative IV) workshop greatly alleviated this problem and allowed us to develop a conditional random fields-based chemical entity recogniser. In order to optimise its performance, we introduced customisations in various aspects of our solution. These include the selection of specialised pre-processing analytics, the incorporation of chemistry knowledge-rich features in the training and application of the statistical model, and the addition of post-processing rules. Our evaluation shows that optimal performance is obtained when our customisations are integrated into the chemical entity recogniser. When its performance is compared with that of state-of-the-art methods, under comparable experimental settings, our solution achieves competitive advantage. We also show that our recogniser that uses a model trained on the CHEMDNER corpus is suitable for recognising names in a wide range of corpora, consistently outperforming two popular chemical NER tools. The contributions resulting from this work are two-fold. Firstly, we present the details of a chemical entity recognition methodology that has demonstrated performance at a competitive, if not superior, level as that of state-of-the-art methods. Secondly, the developed suite of solutions has been made publicly available as a configurable workflow in the interoperable text mining workbench Argo. This allows interested users to conveniently apply and evaluate our solutions in the context of other chemical text mining tasks.