Prediction Modeling Using EHR Data Challenges, Strategies, and a Comparison of Machine Learning Approaches

Prediction Modeling Using EHR Data Challenges, Strategies, and a Comparison of Machine Learning Approaches
复制标题

DOI:
10.1097/mlr.0b013e3181de9e17
复制
发表时间:
2010-06-01
期刊:
影响因子:
3
通讯作者:
Stewart, Walter F.
Stewart, Walter F.
中科院分区:
医学3区
文献类型:
--
作者:
Wu, Jionglin;Roy, Jason;Stewart, Walter F.

文献摘要

被引文献

相似文献

背景:电子健康记录 (EHR) 数据库包含大量有关患者的信息。 Boosting 和支持向量机 (SVM) 等机器学习技术有可能从 EHR 数据中识别患有严重疾病(例如心脏病)的高风险患者。然而,这些技术尚未得到广泛测试。 目的:使用应用于 EHR 数据的机器学习技术,对临床诊断实际日期前 6 个月以上的心力衰竭检测进行建模。比较 Logistic 回归、SVM 和 Boosting 以及各种变量选择方法在心力衰竭预测中的性能。 研究设计:对 Geisinger Clinic 初级保健患者以及 2001 年至 2006 年 EHR 数据中的数据和 2003 年至 2006 年期间诊断为心力衰竭的患者进行了识别。在此巢式病例对照研究中,随机选择性别、年龄和临床相匹配的对照。测量:使用 10 倍交叉验证计算每种方法的受试者操作特征曲线的曲线下面积 (AUC)。比较每种方法选择的变量数量。结果:基于贝叶斯信息准则的模型选择的逻辑回归提供了最简约的模型,平均选择约10个变量,同时保持较高的AUC(10倍交叉验证中为0.77)。具有严格变量重要性阈值的Boosting提供了相似的性能。结论:使用逻辑回归和Boosting在临床诊断前6个月以上预测心力衰竭,AUC约为0.76。即使采用严格的模型选择标准,也能取得这些结果。 SVM 的性能最差,可能是因为数据不平衡。
Background: Electronic health record (EHR) databases contain vast amounts of information about patients. Machine learning techniques such as Boosting and support vector machine (SVM) can potentially identify patients at high risk for serious conditions, such as heart disease, from EHR data. However, these techniques have not yet been widely tested.Objective: To model detection of heart failure more than 6 months before the actual date of clinical diagnosis using machine learning techniques applied to EHR data. To compare the performance of logistic regression, SVM, and Boosting, along with various variable selection methods in heart failure prediction.Research Design: Geisinger Clinic primary care patients with data in the EHR data from 2001 to 2006 diagnosed with heart failure between 2003 and 2006 were identified. Controls were randomly selected matched on sex, age, and clinic for this nested case-control study.Measures: Area under the curve (AUC) of receiver operator characteristic curve was computed for each method using 10-fold cross-validation. The number of variables selected by each method was compared.Results: Logistic regression with model selection based on Bayesian information criterion provided the most parsimonious model, with about 10 variables selected on average, while maintaining a high AUC (0.77 in 10-fold cross-validation). Boosting with strict variable importance threshold provided similar performance.Conclusions: Heart failure was predicted more than 6 months before clinical diagnosis, with AUC of about 0.76, using logistic regression and Boosting. These results were achieved even with strict model selection criteria. SVM had the poorest performance, possibly because of imbalanced data.