Applying active learning to high-throughput phenotyping algorithms for electronic health records data

Applying active learning to high-throughput phenotyping algorithms for electronic health records data
复制标题

DOI:
10.1136/amiajnl-2013-001945
复制
发表时间:
2013-12-01
影响因子:
6.4
通讯作者:
Xu, Hua
Xu, Hua
中科院分区:
管理学2区
文献类型:
--
作者:
Chen, Yukun;Carroll, Robert J.;Xu, Hua

文献摘要

被引文献

相似文献

基于监督机器学习(ML)算法的可推广的高通量表型方法可以显著加快电子健康记录数据在临床和转化研究中的使用。然而,它们通常需要大量带注释的样本,这是昂贵和耗时的审查。我们研究了主动学习(AL)在基于ml的表型算法中的使用。我们将不确定采样AL方法与基于支持向量机的表型算法相结合,并使用三个注释疾病队列(包括类风湿关节炎(RA)、结直肠癌(CRC)和静脉血栓栓塞(VTE))评估其性能。我们使用两种类型的特征集来研究性能:未精炼的特征集,其中至少包含从笔记和账单代码中提取的所有临床概念;以及由领域专家选择的更小的精炼特征集。将人工智能的性能与基于随机抽样的被动学习(PL)方法进行了比较。结果我们的评估显示,AL在三个表型任务上优于PL。当在RA和CRC任务中使用未细化的特征时,人工智能将实现曲线下面积(AUC)得分0.95所需的注释样本数量分别减少了68%和23%。AL还实现了68%的VTE减少,使用精细特征的最佳AUC为0.70。正如预期的那样,改进的特征提高了表型分类器的性能,并且需要更少的注释样本。结论本研究表明,AL可用于基于ml的表型分析方法。此外,人工智能与基于领域知识的特征工程相结合,可以开发出高效、通用的表型方法。
Objectives Generalizable, high-throughput phenotyping methods based on supervised machine learning (ML) algorithms could significantly accelerate the use of electronic health records data for clinical and translational research. However, they often require large numbers of annotated samples, which are costly and time-consuming to review. We investigated the use of active learning (AL) in ML-based phenotyping algorithms.Methods We integrated an uncertainty sampling AL approach with support vector machines-based phenotyping algorithms and evaluated its performance using three annotated disease cohorts including rheumatoid arthritis (RA), colorectal cancer (CRC), and venous thromboembolism (VTE). We investigated performance using two types of feature sets: unrefined features, which contained at least all clinical concepts extracted from notes and billing codes; and a smaller set of refined features selected by domain experts. The performance of the AL was compared with a passive learning (PL) approach based on random sampling.Results Our evaluation showed that AL outperformed PL on three phenotyping tasks. When unrefined features were used in the RA and CRC tasks, AL reduced the number of annotated samples required to achieve an area under the curve (AUC) score of 0.95 by 68% and 23%, respectively. AL also achieved a reduction of 68% for VTE with an optimal AUC of 0.70 using refined features. As expected, refined features improved the performance of phenotyping classifiers and required fewer annotated samples.Conclusions This study demonstrated that AL can be useful in ML-based phenotyping methods. Moreover, AL and feature engineering based on domain knowledge could be combined to develop efficient and generalizable phenotyping methods.