Comparison of Approaches for Heart Failure Case Identification From Electronic Health Record Data.

Comparison of Approaches for Heart Failure Case Identification From Electronic Health Record Data.
复制标题

DOI:
10.1001/jamacardio.2016.3236
复制
发表时间:
2016-12-01
期刊:
影响因子:
24
通讯作者:
Sontag D
Sontag D
中科院分区:
医学1区
文献类型:
--
作者:
Blecker S;Katz SD;Horwitz LI;Kuperman G;Park H;Gold A;Sontag D

文献摘要

被引文献

相似文献

需要准确、实时的病例识别,以便有针对性地进行干预,以提高心力衰竭住院患者的质量和预后。问题列表可能对识别病例很有用,但通常不准确或不完整。机器学习方法可以提高识别的准确性,但会受到实现复杂性的限制。开发算法,使用现成的临床数据来识别住院期间的心力衰竭患者。我们对一家学术医疗中心的住院情况进行了回顾性研究。2013年1月1日之后入院,2015年2月28日之前出院的患者入院时间为18年,≥的住院人数也包括在内。从75%的随机住院样本中,我们开发了五种使用电子健康记录数据识别心力衰竭的算法:1)问题列表上的心力衰竭;2)至少三个特征之一的存在:问题列表上的心力衰竭、住院患者环状利尿剂或脑利钠肽≥500pg/ml;3)30个临床相关的结构化数据元素的Logistic回归;4)使用非结构化笔记的机器学习方法;5)同时使用结构化和非结构化数据的机器学习方法。心力衰竭诊断,基于出院诊断和医生对样本图表的审查。在47119名住院患者中,6549人(13.9%)被诊断为心力衰竭出院。将心力衰竭纳入问题列表(算法1)对心力衰竭识别的灵敏度为0.40,阳性预测值(PPV)为0.96。算法2以0.64的PPV为代价,将灵敏度提高到0.77。算法3、4和5的接收器工作曲线(AUC)下的面积分别为0.953、0.969和0.974。当PPV为0.9时,这些算法的关联灵敏度分别为0.68、0.77和0.83。问题清单不足以实时识别住院的心力衰竭患者。使用自由文本的机器学习的高预测精度表明,在未来的电子病历系统中支持这种分析可以改善队列识别。
Accurate, real-time case identification is needed to target interventions to improve quality and outcomes for hospitalized patients with heart failure. Problem lists may be useful for case identification, but are often inaccurate or incomplete. Machine learning approaches may improve accuracy of identification but can be limited by complexity of implementation. To develop algorithms that use readily available clinical data to identify heart failure patients while in the hospital. We performed a retrospective study of hospitalizations at an academic medical center. Hospitalizations for patients≥18 years who were admitted after January 1, 2013 and discharged prior to February 28, 2015 were included. From a random 75% sample of hospitalizations, we developed five algorithms for heart failure identification using electronic health record (EHR) data: 1) heart failure on problem list; 2) presence of at least one of three characteristics: heart failure on problem list, inpatient loop diuretic, or brain natriuretic peptide≥500 pg/ml; 3) logistic regression of 30 clinically relevant structured data elements; 4) machine learning approach using unstructured notes; 5) machine learning approach using both structured and unstructured data. Heart failure diagnosis, based on discharge diagnosis and physician review of sampled charts. Of 47,119 included hospitalizations, 6,549 (13.9%) had a discharge diagnosis of heart failure. Inclusion of heart failure on the problem list (algorithm 1) had a sensitivity of 0.40 and positive predictive value (PPV) of 0.96 for heart failure identification. Algorithm 2 improved sensitivity to 0.77 at the expense of PPV of 0.64. Algorithms 3, 4, and 5 had areas under the receiver operating curves (AUCs) of 0.953, 0.969, and 0.974, respectively. With PPV of 0.9, these algorithms had associated sensitivities of 0.68, 0.77, and 0.83, respectively. The problem list is insufficient for real-time identification of hospitalized patients with heart failure. The high predictive accuracy of machine learning using free text demonstrates that support of such analytics in future EHR systems can improve cohort identification.