TaggerOne: joint named entity recognition and normalization with semi-Markov Models

TaggerOne: joint named entity recognition and normalization with semi-Markov Models
复制标题

DOI:
10.1093/bioinformatics/btw343
复制
发表时间:
2016-09-15
期刊:
影响因子:
5.8
通讯作者:
Lu, Zhiyong
Lu, Zhiyong
中科院分区:
生物学3区
文献类型:
--
作者:
Leaman, Robert;Lu, Zhiyong

文献摘要

被引文献

相似文献

动机:文本挖掘越来越多地用于管理生物医学文献的加速步伐。许多文本挖掘应用程序依赖于准确的命名实体识别(NER)和规范化(接地)。虽然对于NER存在可针对许多实体类型训练的高性能机器学习方法,但归一化方法通常专用于单个实体类型。NER和规范化系统通常也用于一个串行管道,造成级联错误和限制的NER系统的能力,直接利用所提供的词汇信息的normalization.Methods:我们提出了第一个机器学习模型,联合NER和规范化在训练和预测。该模型可针对任意实体类型进行训练,由半马尔可夫结构线性分类器组成,具有用于NER的丰富特征方法和用于规范化的监督语义索引。我们还介绍了TaggerOne,我们的模型作为一个通用的工具包,联合NER和规范化的Java实现。TaggerOne不特定于任何实体类型,只需要注释的训练数据和相应的词典,并已优化为高throughput.Results:我们验证了TaggerOne与多个黄金标准语料库包含提及和概念级别的注释。基准测试结果表明,TaggerOne在疾病(NCBI疾病语料库,NER f-分数:0.829,标准化f-分数:0.807)和化学品(BioCreative 5 CDR语料库,NER f-分数:0.914,标准化f-分数0.895)上实现了高性能。这些结果与现有技术相比是有利的,尽管模型具有更大的灵活性。我们的结论是,联合建模NER和规范化大大提高了性能。
Motivation: Text mining is increasingly used to manage the accelerating pace of the biomedical literature. Many text mining applications depend on accurate named entity recognition (NER) and normalization (grounding). While high performing machine learning methods trainable for many entity types exist for NER, normalization methods are usually specialized to a single entity type. NER and normalization systems are also typically used in a serial pipeline, causing cascading errors and limiting the ability of the NER system to directly exploit the lexical information provided by the normalization.Methods: We propose the first machine learning model for joint NER and normalization during both training and prediction. The model is trainable for arbitrary entity types and consists of a semi-Markov structured linear classifier, with a rich feature approach for NER and supervised semantic indexing for normalization. We also introduce TaggerOne, a Java implementation of our model as a general toolkit for joint NER and normalization. TaggerOne is not specific to any entity type, requiring only annotated training data and a corresponding lexicon, and has been optimized for high throughput.Results: We validated TaggerOne with multiple gold-standard corpora containing both mention- and concept-level annotations. Benchmarking results show that TaggerOne achieves high performance on diseases (NCBI Disease corpus, NER f-score: 0.829, normalization f-score: 0.807) and chemicals (BioCreative 5 CDR corpus, NER f-score: 0.914, normalization f-score 0.895). These results compare favorably to the previous state of the art, notwithstanding the greater flexibility of the model. We conclude that jointly modeling NER and normalization greatly improves performance.