Strategies for handling missing data in electronic health record derived data.

Strategies for handling missing data in electronic health record derived data.
复制标题

DOI:
10.13063/2327-9214.1035
复制
发表时间:
2013
期刊:
EGEMS (Washington, DC)
影响因子:
--
通讯作者:
Kattan MW
Kattan MW
中科院分区:
其他
文献类型:
--
作者:
Wells BJ;Chagin KM;Nowacki AS;Kattan MW

文献摘要

被引文献

相似文献

电子健康记录(EHR)提供了丰富的数据,这些数据对于改善以患者为中心的结果至关重要,尽管这些数据可能会带来重大的统计挑战。特别是,电子病历数据包含大量缺失的信息,如果不加以处理,可能会降低所得出结论的有效性。由于有时很难区分丢失数据和负值,因此很难正确处理EHR数据中的丢失数据问题。例如,没有记录的心力衰竭病史的患者可能真的没有疾病,或者临床医生可能只是没有记录这种情况。减少EHR系统中丢失数据的方法来自多个角度,包括:增加结构化数据文档,减少数据输入错误,以及利用文本解析/自然语言处理。本文主要研究处理缺失数据的分析方法,主要是多重补偿。典型的电子病历系统中可用的变量范围广泛,为缓解数据缺失造成的潜在偏差提供了丰富的信息。数据丢失的概率可能与疾病严重性和医疗保健利用率有关,因为不健康的患者更有可能患有并存疾病,而每次与医疗保健系统的互动都提供了记录的机会。因此,任何归罪例程都应该包括评估总体健康状况(例如Charlson共病指数)和医疗保健利用(例如遇到次数)的预测变量,即使这些并发症和患者遇到的情况与感兴趣的疾病无关。将电子健康记录数据与其他信息来源(如国家死亡指数和人口普查数据)联系起来,也可以提供较少偏见的变量来进行推算。有必要利用EHR数据进行额外的方法学研究,并改进临床研究人员的流行病学培训。
Electronic health records (EHRs) present a wealth of data that are vital for improving patient-centered outcomes, although the data can present significant statistical challenges. In particular, EHR data contains substantial missing information that if left unaddressed could reduce the validity of conclusions drawn. Properly addressing the missing data issue in EHR data is complicated by the fact that it is sometimes difficult to differentiate between missing data and a negative value. For example, a patient without a documented history of heart failure may truly not have disease or the clinician may have simply not documented the condition. Approaches for reducing missing data in EHR systems come from multiple angles, including: increasing structured data documentation, reducing data input errors, and utilization of text parsing / natural language processing. This paper focuses on the analytical approaches for handling missing data, primarily multiple imputation. The broad range of variables available in typical EHR systems provide a wealth of information for mitigating potential biases caused by missing data. The probability of missing data may be linked to disease severity and healthcare utilization since unhealthier patients are more likely to have comorbidities and each interaction with the health care system provides an opportunity for documentation. Therefore, any imputation routine should include predictor variables that assess overall health status (e.g. Charlson Comorbidity Index) and healthcare utilization (e.g. number of encounters) even when these comorbidities and patient encounters are unrelated to the disease of interest. Linking the EHR data with other sources of information (e.g. National Death Index and census data) can also provide less biased variables for imputation. Additional methodological research with EHR data and improved epidemiological training of clinical investigators is warranted.