Automated feature selection of predictors in electronic medical records data

Automated feature selection of predictors in electronic medical records data
复制标题

DOI:
10.1111/biom.12987
复制
发表时间:
2019-03-01
期刊:
影响因子:
1.9
通讯作者:
Cai, Tianxi
Cai, Tianxi
中科院分区:
数学3区
文献类型:
--
作者:
Gronsbell, Jessica;Minnier, Jessica;Cai, Tianxi

文献摘要

被引文献

相似文献

由于难以提取准确的疾病表型数据,将电子健康记录(EHR)用于转化研究可能具有挑战性。从历史上看,用于注释表型的EHR算法要么是基于规则的,要么是通过劳动密集型病历审查管理的计费代码和黄金标准标签进行训练的。这些过于简单的算法往往具有不可预测的跨机构的可移植性,并且由于不精确的计费代码,对于许多疾病表型具有低准确性。最近,已经开发了更复杂的机器学习算法来提高EHR表型分析算法的鲁棒性和准确性。这些算法通常通过监督学习进行训练,将金标准标签与广泛的候选特征相关联,包括账单代码,程序代码,药物处方和通过自然语言处理(NLP)从叙述性笔记中提取的相关临床概念。然而,由于黄金标准标记的时间密集性,训练集的大小往往不足以构建具有大量候选特征的EHR的可推广算法。为了减少候选预测因子的数量,进而提高模型性能,我们提出了一种完全基于未标记观测的自动特征选择方法。所提出的方法生成一个全面的替代的基础表型与疾病状态的无监督聚类的基础上,几个高度预测的功能,如诊断代码和提到的疾病在整个EHR数据集的文本字段。然后用估计的结果和剩余的协变量构建稀疏回归模型,以识别那些最能提供感兴趣的表型信息的特征。依靠Li和Duan(1989)的结果,我们证明了可以通过拟合基于替代模型来实现潜在表型模型的变量选择。我们探讨了我们的方法在数值模拟中的性能,并提出了建立在合作伙伴健康系统的大型EHR数据集市上的风湿性关节炎(RA)预测模型的结果,该数据集市由计费代码和NLP术语组成。实证结果表明,我们的程序减少了表型分析所需的黄金标准标签的数量,从而利用EHR数据的自动化功能并提高效率。
The use of Electronic Health Records (EHR) for translational research can be challenging due to difficulty in extracting accurate disease phenotype data. Historically, EHR algorithms for annotating phenotypes have been either rule-based or trained with billing codes and gold standard labels curated via labor intensive medical chart review. These simplistic algorithms tend to have unpredictable portability across institutions and low accuracy for many disease phenotypes due to imprecise billing codes. Recently, more sophisticated machine learning algorithms have been developed to improve the robustness and accuracy of EHR phenotyping algorithms. These algorithms are typically trained via supervised learning, relating gold standard labels to a wide range of candidate features including billing codes, procedure codes, medication prescriptions and relevant clinical concepts extracted from narrative notes via Natural Language Processing (NLP). However, due to the time intensiveness of gold standard labeling, the size of the training set is often insufficient to build a generalizable algorithm with the large number of candidate features extracted from EHR. To reduce the number of candidate predictors and in turn improve model performance, we present an automated feature selection method based entirely on unlabeled observations. The proposed method generates a comprehensive surrogate for the underlying phenotype with an unsupervised clustering of disease status based on several highly predictive features such as diagnosis codes and mentions of the disease in text fields available in the entire set of EHR data. A sparse regression model is then built with the estimated outcomes and remaining covariates to identify those features most informative of the phenotype of interest. Relying on the results of Li and Duan (1989), we demonstrate that variable selection for the underlying phenotype model can be achieved by fitting the surrogate-based model. We explore the performance of our methods in numerical simulations and present the results of a prediction model for Rheumatoid Arthritis (RA) built on a large EHR data mart from the Partners Health System consisting of billing codes and NLP terms. Empirical results suggest that our procedure reduces the number of gold-standard labels necessary for phenotyping thereby harnessing the automated power of EHR data and improving efficiency.