Scalable and accurate deep learning with electronic health records

Scalable and accurate deep learning with electronic health records
复制标题

DOI:
10.1038/s41746-018-0029-1
复制
发表时间:
2018-05-08
影响因子:
15.2
通讯作者:
Dean, Jeffrey
Dean, Jeffrey
中科院分区:
医学1区
文献类型:
--
作者:
Rajkomar, Alvin;Oren, Eyal;Dean, Jeffrey

文献摘要

被引文献

相似文献

利用电子健康记录 (EHR) 数据进行预测建模预计将推动个性化医疗并提高医疗保健质量。构建预测统计模型通常需要从标准化 EHR 数据中提取精心策划的预测变量,这是一个劳动密集型过程,会丢弃每个患者记录中的绝大多数信息。我们建议基于快速医疗互操作性资源 (FHIR) 格式来表示患者的整个原始 EHR 记录。我们证明,使用这种表示的深度学习方法能够准确预测来自多个中心的多个医疗事件,而无需特定地点的数据协调。我们使用来自两个美国学术医疗中心的去识别化 EHR 数据验证了我们的方法,其中包括 216,221 名住院至少 24 小时的成年患者。在我们提出的顺序格式中,这批 EHR 数据总共展开为 46,864,534,945 个数据点,包括临床记录。深度学习模型在以下任务上实现了高精度预测:院内死亡率(接受者操作曲线下面积 [AUROC] 跨站点 0.93-0.94)、30 天计划外再入院 (AUROC 0.75-0.76)、延长住院时间 (AUROC 0.85-0.86) 以及患者的所有最终出院诊断(频率加权 AUROC) 0.90)。这些模型在所有情况下都优于传统的临床使用的预测模型。我们相信这种方法可用于为各种临床场景创建准确且可扩展的预测。在特定预测的案例研究中,我们证明神经网络可用于从患者图表中识别相关信息。
Predictive modeling with electronic health record (EHR) data is anticipated to drive personalized medicine and improve healthcare quality. Constructing predictive statistical models typically requires extraction of curated predictor variables from normalized EHR data, a labor-intensive process that discards the vast majority of information in each patient's record. We propose a representation of patients' entire raw EHR records based on the Fast Healthcare Interoperability Resources (FHIR) format. We demonstrate that deep learning methods using this representation are capable of accurately predicting multiple medical events from multiple centers without site-specific data harmonization. We validated our approach using de-identified EHR data from two US academic medical centers with 216,221 adult patients hospitalized for at least 24 h. In the sequential format we propose, this volume of EHR data unrolled into a total of 46,864,534,945 data points, including clinical notes. Deep learning models achieved high accuracy for tasks such as predicting: in-hospital mortality (area under the receiver operator curve [AUROC] across sites 0.93-0.94), 30-day unplanned readmission (AUROC 0.75-0.76), prolonged length of stay (AUROC 0.85-0.86), and all of a patient's final discharge diagnoses (frequency-weighted AUROC 0.90). These models outperformed traditional, clinically-used predictive models in all cases. We believe that this approach can be used to create accurate and scalable predictions for a variety of clinical scenarios. In a case study of a particular prediction, we demonstrate that neural networks can be used to identify relevant information from the patient's chart.