Sparse data bias: a problem hiding in plain sight

Sparse data bias: a problem hiding in plain sight
复制标题

DOI:
10.1136/bmj.i1981
复制
发表时间:
2016-04-27
影响因子:
105.7
通讯作者:
Altman, Douglas G.
Altman, Douglas G.
中科院分区:
医学1区
文献类型:
--
作者:
Greenland, Sander;Mansournia, Mohammad Ali;Altman, Douglas G.

文献摘要

被引文献

相似文献

治疗或其他暴露对结局事件的影响通常通过风险比、发生率或比值来衡量。这些指标的调整版本通常通过最大似然回归(例如,logistic、Poisson或考克斯模型)进行估计。但是,当数据缺乏足够的病例数来描述暴露水平和结果水平的组合时,对效应测量的估计可能会有严重的偏差。这种偏差甚至可以发生在相当大的数据集,因此通常被称为稀疏数据偏差。通过对潜在混杂变量进行回归调整,偏差可能会出现或恶化;在极端情况下,所得到的估计值可能是不可能的巨大值,甚至是无限值,这些值是数据稀疏的无意义伪影。根据背景资料,这种估计通货膨胀可能是显而易见的,但很少被注意到,更不用说在研究报告中解释了。我们概述了简单的方法,用于检测和处理的问题,特别是集中在惩罚估计,这可以很容易地执行与常见的软件包。
Effects of treatment or other exposure on outcome events are commonly measured by ratios of risks, rates, or odds. Adjusted versions of these measures are usually estimated by maximum likelihood regression (eg, logistic, Poisson, or Cox modelling). But resulting estimates of effect measures can have serious bias when the data lack adequate case numbers for some combination of exposure and outcome levels. This bias can occur even in quite large datasets and is hence often termed sparse data bias. The bias can arise or be worsened by regression adjustment for potentially confounding variables; in the extreme, the resulting estimates could be impossibly huge or even infinite values that are meaningless artefacts of data sparsity. Such estimate inflation might be obvious in light of background information, but is rarely noted let alone accounted for in research reports. We outline simple methods for detecting and dealing with the problem focusing especially on penalised estimation, which can be easily performed with common software packages.