A General Framework for Considering Selection Bias in EHR-Based Studies: What Data Are Observed and Why?

A General Framework for Considering Selection Bias in EHR-Based Studies: What Data Are Observed and Why?
复制标题

DOI:
10.13063/2327-9214.1203
复制
发表时间:
2016-01-01
期刊:
EGEMS (Washington, DC)
影响因子:
--
通讯作者:
Daniels, Michael
Daniels, Michael
中科院分区:
其他
文献类型:
--
作者:
Haneuse, Sebastien;Daniels, Michael

文献摘要

被引文献

相似文献

电子健康记录 (EHR) 数据越来越被视为具有成本效益的比较效果研究 (CER) 的资源。由于收集 EHR 数据主要用于临床和/或计费目的,因此将其用于 CER 需要考虑许多方法学挑战,包括由于缺乏随机性而可能出现混杂偏倚,以及由于缺失数据而导致选择偏倚。与最近关于基于 EHR 的 CER 中混杂偏差的文献相比,实际上没有关注选择偏差,这可能是因为人们相信可以轻松应用缺失数据的标准方法。然而,此类方法依赖于对可用/缺失 EHR 数据过于简单化的看法,因此它们在 EHR 设置中的应用通常无法完全控制选择偏差。受我们在一项正在进行的基于 EHR 的抗抑郁治疗选择和长期体重变化比较有效性研究中面临的挑战的启发,我们提出了一个新的基于 EHR 的 CER 选择偏差的通用框架。至关重要的是,该框架提供了一个结构,研究人员可以在该结构中考虑患者和医疗保健提供者做出的众多决策之间复杂的相互作用,这些决策导致 EHR 系统中记录的健康相关信息,以及 EHR 系统本身的广泛差异。这反过来又提供了以下结构:(i)可以增强有关缺失数据的假设的透明度,(ii)可以得出与每个决策相关的因素,以及(iii)统计方法可以更好地与数据的复杂性保持一致。
Electronic health records (EHR) data are increasingly seen as a resource for cost-effective comparative effectiveness research (CER). Since EHR data are collected primarily for clinical and/or billing purposes, their use for CER requires consideration of numerous methodologic challenges including the potential for confounding bias, due to a lack of randomization, and for selection bias, due to missing data. In contrast to the recent literature on confounding bias in EHR-based CER, virtually no attention has been paid to selection bias possibly due to the belief that standard methods for missing data can be readily-applied. Such methods, however, hinge on an overly simplistic view of the available/missing EHR data, so that their application in the EHR setting will often fail to completely control selection bias. Motivated by challenges we face in an on-going EHR-based comparative effectiveness study of choice of antidepressant treatment and long-term weight change, we propose a new general framework for selection bias in EHR-based CER. Crucially, the framework provides structure within which researchers can consider the complex interplay between numerous decisions, made by patients and health care providers, which give rise to health-related information being recorded in the EHR system, as well as the wide variability across EHR systems themselves. This, in turn, provides structure within which: (i) the transparency of assumptions regarding missing data can be enhanced, (ii) factors relevant to each decision can be elicited, and (iii) statistical methods can be better aligned with the complexity of the data.