Applying active learning to assertion classification of concepts in clinical text.

Applying active learning to assertion classification of concepts in clinical text.
复制标题

DOI:
10.1016/j.jbi.2011.11.003
复制
发表时间:
2012-04
影响因子:
4.5
通讯作者:
Xu, Hua
Xu, Hua
中科院分区:
医学3区
文献类型:
--
作者:
Chen, Yukun;Mani, Subramani;Xu, Hua

文献摘要

参考文献

被引文献

相似文献

用于临床自然语言处理(NLP)研究的监督机器学习方法需要大量带注释的样本,由于医生的参与,构建这些样本非常昂贵。主动学习是一种从大样本库中主动采样的方法,它提供了一种替代解决方案。它在分类中的主要目标是减少注释工作,同时保持预测模型的质量。然而,很少有研究调查其在临床NLP中的用途。本文报告了一个应用主动学习的临床文本分类任务:确定临床概念的断言状态。本研究使用了2010年i2 b2/VA临床NLP挑战赛中断言分类任务的注释语料库。我们实现了几个现有的和新开发的主动学习算法,并评估其使用。根据AUC(曲线下面积)评分的平均学习曲线下面积,在总体ALC评分中报告结局。结果表明,当使用相同数量的注释样本时,主动学习策略可以产生更好的分类模型(最佳ALC - 0.7715)比被动学习方法(随机抽样)(ALC - 0.7411)。此外,为了达到相同的分类性能,主动学习策略需要的样本比随机抽样方法少。例如,为了实现0.79的AUC,随机采样方法使用了32个样本,而我们最好的主动学习算法只需要12个样本,减少了62.5%的手动注释工作。
Supervised machine learning methods for clinical natural language processing (NLP) research require a large number of annotated samples, which are very expensive to build because of the involvement of physicians. Active learning, an approach that actively samples from a large pool, provides an alternative solution. Its major goal in classification is to reduce the annotation effort while maintaining the quality of the predictive model. However, few studies have investigated its uses in clinical NLP. This paper reports an application of active learning to a clinical text classification task: to determine the assertion status of clinical concepts. The annotated corpus for the assertion classification task in the 2010 i2b2/VA Clinical NLP Challenge was used in this study. We implemented several existing and newly developed active learning algorithms and assessed their uses. The outcome is reported in the global ALC score, based on the Area under the average Learning Curve of the AUC (Area Under the Curve) score. Results showed that when the same number of annotated samples was used, active learning strategies could generate better classification models (best ALC – 0.7715) than the passive learning method (random sampling) (ALC – 0.7411). Moreover, to achieve the same classification performance, active learning strategies required fewer samples than the random sampling method. For example, to achieve an AUC of 0.79, the random sampling method used 32 samples, while our best active learning algorithm required only 12 samples, a reduction of 62.5% in manual annotation effort.
DOI: 10.1136/jamia.1994.95236146
发表时间: 1994-03-01
影响因子: 6.4
作者:
FRIEDMAN, C;ALDERSON, PO;JOHNSON, SB
通讯作者: JOHNSON, SB
DOI: 10.1093/jee/39.2.269
发表时间: 1945-01-01
期刊: BIOMETRICS BULLETIN
影响因子: --
作者:
WILCOXON, F
通讯作者: WILCOXON, F
DOI: 10.1021/ci049810a
发表时间: 2004-11-01
期刊: JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子: --
作者:
Liu, Y
通讯作者: Liu, Y
DOI: 10.1136/jamia.2010.004200
发表时间: 2010-09-01
影响因子: 6.4
作者:
Uzuner, Oezlem;Solti, Imre;Cadag, Eithon
通讯作者: Cadag, Eithon
DOI: 10.1162/153244302760185243
发表时间: 2002-12-01
影响因子: 6
作者:
Tong, S;Koller, D
通讯作者: Koller, D