tmChem: a high performance approach for chemical named entity recognition and normalization.

tmChem: a high performance approach for chemical named entity recognition and normalization.
复制标题

DOI:
10.1186/1758-2946-7-s1-s3
复制
发表时间:
2015
影响因子:
8.6
通讯作者:
Lu Z
Lu Z
中科院分区:
化学2区
文献类型:
--
作者:
Leaman R;Wei CH;Lu Z

文献摘要

被引文献

相似文献

化合物和药物是生物医学研究中的一类重要实体,在包括临床医学在内的广泛应用中具有巨大潜力。在文献中查找化学命名实体是化学文本挖掘管道中的一个有用步骤,用于识别文献中讨论的化学提及、其属性及其关系。我们介绍了 tmChem 系统,这是一种化学命名实体识别器,通过将两个独立的机器学习模型组合在一个集合中而创建。我们使用最近 CHEMDNER 任务的一部分发布的语料库来开发和评估 tmChem,在 CEM 子任务(提及级评估)上实现了 0.8739 的微平均 f 测量,在 CDI 子任务(抽象级评估)上实现了 0.8745 f 测量。我们还报告了高召回率组合(CEM 为 0.9212,CDI 为 0.9224)。 tmChem 在 CEM 子任务的 CHEMDNER 任务中实现了最高的 f 测量,并且高召回率变体在 CEM 和 CDI 任务上都实现了最高的召回率。我们报告说,tmChem 是最先进的化学命名实体识别工具,并且化学命名实体识别的性能现已追平(或超过)之前报告的基因和疾病的性能。未来的研究应侧重于命名实体识别和规范化步骤之间更紧密的集成,以提高性能。 tmChem 两种模型的源代码和训练模型可从以下网址获取:http://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/tmChem。在 PubMed 上运行 tmChem(模型 2)的结果可在 PubTator 中获取:http://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/PubTator
Chemical compounds and drugs are an important class of entities in biomedical research with great potential in a wide range of applications, including clinical medicine. Locating chemical named entities in the literature is a useful step in chemical text mining pipelines for identifying the chemical mentions, their properties, and their relationships as discussed in the literature. We introduce the tmChem system, a chemical named entity recognizer created by combining two independent machine learning models in an ensemble. We use the corpus released as part of the recent CHEMDNER task to develop and evaluate tmChem, achieving a micro-averaged f-measure of 0.8739 on the CEM subtask (mention-level evaluation) and 0.8745 f-measure on the CDI subtask (abstract-level evaluation). We also report a high-recall combination (0.9212 for CEM and 0.9224 for CDI). tmChem achieved the highest f-measure reported in the CHEMDNER task for the CEM subtask, and the high recall variant achieved the highest recall on both the CEM and CDI tasks. We report that tmChem is a state-of-the-art tool for chemical named entity recognition and that performance for chemical named entity recognition has now tied (or exceeded) the performance previously reported for genes and diseases. Future research should focus on tighter integration between the named entity recognition and normalization steps for improved performance. The source code and a trained model for both models of tmChem is available at: http://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/tmChem. The results of running tmChem (Model 2) on PubMed are available in PubTator: http://www.ncbi.nlm.nih.gov/CBBresearch/Lu/Demo/PubTator