A study of active learning methods for named entity recognition in clinical text.

A study of active learning methods for named entity recognition in clinical text.
复制标题

DOI:
10.1016/j.jbi.2015.09.010
复制
发表时间:
2015-12
影响因子:
4.5
通讯作者:
Xu H
Xu H
中科院分区:
医学3区
文献类型:
--
作者:
Chen Y;Lasko TA;Mei Q;Denny JC;Xu H

文献摘要

被引文献

相似文献

命名实体识别(NER)是一种顺序标注任务,是构建临床自然语言处理(NLP)系统的基本任务之一。基于机器学习(ML)的方法可以实现良好的性能,但它们通常需要大量的注释样本,由于注释领域专家的要求,这些注释样本的构建成本很高。主动学习(AL)是一种与监督ML集成的样本选择方法,旨在最大限度地降低标注成本,同时最大限度地提高基于ML的模型的性能。在这项研究中,我们的目标是开发和评估现有的和新的AL方法的临床NER任务,以确定概念的医疗问题,治疗和实验室测试的临床笔记。使用来自2010年i2 b2/VA NLP挑战赛的注释NER语料库,其中包含349个临床文档和20,423个独特的句子,我们使用三个不同类别的一些现有和新颖的算法模拟AL实验,包括基于不确定性,基于多样性和基线采样策略。他们与使用随机抽样的被动学习进行了比较。生成绘制NER模型的性能相对于估计的注释成本(基于训练集中的句子或单词的数量)的学习曲线,以评估不同的主动学习和被动学习方法,并计算学习曲线下的面积(ALC)得分。基于F-测度与句子数量的学习曲线,不确定性抽样算法优于ALC中的所有其他方法。在ALC中,大多数基于多样性的方法也比随机抽样更好。为了达到0.80的F-测量,与随机抽样相比,基于不确定性抽样的最佳方法可以节省66%的句子注释。对于F-测量与字数的学习曲线,不确定性采样方法再次优于ALC中的所有其他方法。为了实现0.80的F-测量,与随机抽样相比,基于最佳不确定性的方法节省了42%的文字注释。但是最好的基于多样性的方法只减少了7%的注释工作。在模拟环境中,AL方法,特别是基于不确定性采样的方法,似乎显着节省注释成本的临床NER任务。主动学习在临床NER中的实际益处应在实时环境中进一步评估。
Named entity recognition (NER), a sequential labeling task, is one of the fundamental tasks for building clinical natural language processing (NLP) systems. Machine learning (ML) based approaches can achieve good performance, but they often require large amounts of annotated samples, which are expensive to build due to the requirement of domain experts in annotation. Active learning (AL), a sample selection approach integrated with supervised ML, aims to minimize the annotation cost while maximizing the performance of ML-based models. In this study, our goal was to develop and evaluate both existing and new AL methods for a clinical NER task to identify concepts of medical problems, treatments, and lab tests from the clinical notes. Using the annotated NER corpus from the 2010 i2b2/VA NLP challenge that contained 349 clinical documents with 20,423 unique sentences, we simulated AL experiments using a number of existing and novel algorithms in three different categories including uncertainty-based, diversity-based, and baseline sampling strategies. They were compared with the passive learning that uses random sampling. Learning curves that plot performance of the NER model against the estimated annotation cost (based on number of sentences or words in the training set) were generated to evaluate different active learning and the passive learning methods and the area under the learning curve (ALC) score was computed. Based on the learning curves of F-measure vs. number of sentences, uncertainty sampling algorithms outperformed all other methods in ALC. Most diversity-based methods also performed better than random sampling in ALC. To achieve an F-measure of 0.80, the best method based on uncertainty sampling could save 66% annotations in sentences, as compared to random sampling. For the learning curves of F-measure vs. number of words, uncertainty sampling methods again outperformed all other methods in ALC. To achieve 0.80 in F-measure, in comparison to random sampling, the best uncertainty based method saved 42% annotations in words. But the best diversity based method reduced only 7% annotation effort. In the simulated setting, AL methods, particularly uncertainty-sampling based approaches, seemed to significantly save annotation cost for the clinical NER task. The actual benefit of active learning in clinical NER should be further evaluated in a real-time setting.