Development and validation of machine learning models to identify high-risk surgical patients using automatically curated electronic health record data (Pythia): A retrospective, single-site study.

Development and validation of machine learning models to identify high-risk surgical patients using automatically curated electronic health record data (Pythia): A retrospective, single-site study.
复制标题

DOI:
10.1371/journal.pmed.1002701
复制
发表时间:
2018-11
期刊:
影响因子:
15.8
通讯作者:
Sendak M
Sendak M
中科院分区:
医学1区
文献类型:
--
作者:
Corey KM;Kashyap S;Lorenzi E;Lagoo-Deenadayalan SA;Heller K;Whalen K;Balu S;Heflin MT;McDonald SR;Swaminathan M;Sendak M

文献摘要

被引文献

相似文献

Pythia是一个自动化的,临床策划的手术数据管道和存储库,存储来自大型,四元,多站点健康研究所的所有手术患者电子健康记录(EHR)数据,用于数据科学计划。为了更好地从复杂数据中识别高风险手术患者,建立了一个在Pythia上训练的机器学习项目,以预测术后并发症风险。使用自动化的SQL和R代码创建了手术结果的策划数据存储库,该代码从EHR中提取和处理了3700万次临床就诊的患者临床和手术数据。总共194个临床特征,包括患者人口统计数据(例如,年龄、性别、种族)、吸烟状况、药物、合并症、手术信息和手术复杂性的替代物。为了预测术后并发症,对2014年1月1日至2017年1月31日期间接受99,755次侵入性手术的66,370例患者进行了进一步研究。该队列的平均并发症发生率和术后30天死亡率分别为16.0%和0.51%。最小绝对收缩和选择算子(lasso)惩罚逻辑回归,随机森林模型和极端梯度提升决策树在该手术队列中进行了训练,并对14个特定的术后结局分组进行了交叉验证。所得模型的受试者工作特征曲线下面积(AUC)值范围为0.747至0.924,根据最近5个月数据的样本外测试集计算。Lasso惩罚回归被认为是一种高性能模型,提供了临床可解释的可操作见解。最高和最低性能lasso模型预测术后休克和泌尿生殖系统结局的AUC分别为0.924(95% CI:0.901,0.946)和0.780(95% CI:0.752,0.810)。创建了一个需要输入9个数据字段的计算器,以生成14组术后结局的风险评估。确定高风险阈值(任何并发症的风险15%)以识别高风险手术患者。模型的敏感性为76%,特异性为76%。与由临床专家开发的识别高风险患者的诊断方法和ACS NSQIP计算器相比,该工具表现优越,为临床医生估计患者术后风险提供了改进的方法。本研究的局限性包括删除用于分析的数据的缺失。为了机器学习的目的,提取和管理大型本地机构的EHR数据,产生了具有强大预测性能的模型。这些模型可以在临床环境中作为决策支持工具,用于识别高风险患者以及患者评估和护理管理。有必要开展进一步的工作,以评估Pythia风险计算器在临床工作流程中对术后结局的影响,并为未来的机器学习工作优化此数据流。Kristin Corey及其同事利用单站点手术EHR数据管道和存储库,在他们的机构中提出了一种基于机器学习的高风险手术患者检测。大多数术后并发症风险预测模型使用美国外科医师学会(ACS)国家外科质量改进计划(NSQIP)数据。很少有已发表的使用电子健康记录(EHR)数据的术后风险预测模型存在。创建和更新手动数据集(如ACS NSQIP)是时间、人力和成本密集型流程。创建了一个术后结局的策划数据库,将EHR中3700万次临床就诊的患者临床和手术数据提取并处理为194个临床特征。机器学习模型建立在这个数据集上,以预测术后并发症的风险。模型能够以高灵敏度和特异性对术后并发症高危患者进行分类。创建了一个需要输入9个数据字段的在线计算器,以在临床环境中进行风险评估。基于自动提取和策划的EHR手术数据集构建的机器学习模型在检测手术并发症方面具有很强的预测性能。更新该手术数据集的长期成本低于手动更新数据集的成本。
Pythia is an automated, clinically curated surgical data pipeline and repository housing all surgical patient electronic health record (EHR) data from a large, quaternary, multisite health institute for data science initiatives. In an effort to better identify high-risk surgical patients from complex data, a machine learning project trained on Pythia was built to predict postoperative complication risk. A curated data repository of surgical outcomes was created using automated SQL and R code that extracted and processed patient clinical and surgical data across 37 million clinical encounters from the EHRs. A total of 194 clinical features including patient demographics (e.g., age, sex, race), smoking status, medications, comorbidities, procedure information, and proxies for surgical complexity were constructed and aggregated. A cohort of 66,370 patients that had undergone 99,755 invasive procedural encounters between January 1, 2014, and January 31, 2017, was studied further for the purpose of predicting postoperative complications. The average complication and 30-day postoperative mortality rates of this cohort were 16.0% and 0.51%, respectively. Least absolute shrinkage and selection operator (lasso) penalized logistic regression, random forest models, and extreme gradient boosted decision trees were trained on this surgical cohort with cross-validation on 14 specific postoperative outcome groupings. Resulting models had area under the receiver operator characteristic curve (AUC) values ranging between 0.747 and 0.924, calculated on an out-of-sample test set from the last 5 months of data. Lasso penalized regression was identified as a high-performing model, providing clinically interpretable actionable insights. Highest and lowest performing lasso models predicted postoperative shock and genitourinary outcomes with AUCs of 0.924 (95% CI: 0.901, 0.946) and 0.780 (95% CI: 0.752, 0.810), respectively. A calculator requiring input of 9 data fields was created to produce a risk assessment for the 14 groupings of postoperative outcomes. A high-risk threshold (15% risk of any complication) was determined to identify high-risk surgical patients. The model sensitivity was 76%, with a specificity of 76%. Compared to heuristics that identify high-risk patients developed by clinical experts and the ACS NSQIP calculator, this tool performed superiorly, providing an improved approach for clinicians to estimate postoperative risk for patients. Limitations of this study include the missingness of data that were removed for analysis. Extracting and curating a large, local institution’s EHR data for machine learning purposes resulted in models with strong predictive performance. These models can be used in clinical settings as decision support tools for identification of high-risk patients as well as patient evaluation and care management. Further work is necessary to evaluate the impact of the Pythia risk calculator within the clinical workflow on postoperative outcomes and to optimize this data flow for future machine learning efforts. Leveraging a single-site surgical EHR data pipeline and repository, Kristin Corey and colleagues present a machine learning-based detection of high-risk surgical patients at their institution. Most postoperative complication risk prediction models use American College of Surgeons (ACS) National Surgical Quality Improvement Program (NSQIP) data. Few published postoperative risk prediction models using electronic health record (EHR) data exist. Creating and updating manual datasets such as ACS NSQIP are intensive processes with regards to time, labor, and cost. A curated data repository of postoperative outcomes was created that extracted and processed patient clinical and surgical data across 37 million clinical encounters in EHRs into 194 clinical features. Machine learning models were built off this dataset to predict risk of postoperative complications. Models were able to classify patients at high risk of postoperative complication with high sensitivity and specificity. An online calculator requiring input of 9 data fields was created to produce a risk assessment within the clinic environment. Machine leaning models built off an automatically extracted and curated EHR surgical dataset have strong predictive performance for detecting surgical complications. The long-term cost of updating this surgical dataset is lower than that of manually updated datasets.