An active learning-enabled annotation system for clinical named entity recognition

An active learning-enabled annotation system for clinical named entity recognition
复制标题

DOI:
10.1186/s12911-017-0466-9
复制
发表时间:
2017-07-05
影响因子:
3.5
通讯作者:
Xu, Hua
Xu, Hua
中科院分区:
医学3区
文献类型:
--
作者:
Chen, Yukun;Lask, Thomas A.;Xu, Hua

文献摘要

被引文献

相似文献

背景:在构建统计自然语言处理(NLP)模型中,主动学习(AL)在最小化标注成本的同时最大化性能方面显示出了良好的潜力。然而,很少有研究在医学领域的现实环境中调查人工智能。方法:在本研究中,我们利用一种新颖的人工智能算法开发了第一个用于临床命名实体识别(NER)的人工智能注释系统。除了模拟研究来评估新的人工智能算法外,我们还对两名使用该系统的护士进行了用户研究,以评估人工智能在构建临床NER模型的真实世界注释过程中的性能。结果:仿真结果表明,该算法优于传统的人工智能算法和随机抽样算法。然而,用户研究告诉我们一个不同的故事,即对于不同的用户,人工智能方法并不总是比随机抽样更好。结论:我们发现,主动选择的句子所增加的信息含量会被标注时间的增加所抵消。此外,在查询算法中没有考虑标注时间。我们未来的工作包括开发更好的人工智能算法,估计注释时间,并在更大的用户数量下评估系统。
Background: Active learning (AL) has shown the promising potential to minimize the annotation cost while maximizing the performance in building statistical natural language processing (NLP) models. However, very few studies have investigated AL in a real-life setting in medical domain.Methods: In this study, we developed the first AL-enabled annotation system for clinical named entity recognition (NER) with a novel AL algorithm. Besides the simulation study to evaluate the novel AL algorithm, we further conducted user studies with two nurses using this system to assess the performance of AL in real world annotation processes for building clinical NER models.Results: The simulation results show that the novel AL algorithm outperformed traditional AL algorithm and random sampling. However, the user study tells a different story that AL methods did not always perform better than random sampling for different users.Conclusions: We found that the increased information content of actively selected sentences is strongly offset by the increased time required to annotate them. Moreover, the annotation time was not considered in the querying algorithms. Our future work includes developing better AL algorithms with the estimation of annotation time and evaluating the system with larger number of users.