Local model uncertainty and incomplete-data bias

Local model uncertainty and incomplete-data bias
复制标题

DOI:
10.1111/j.1467-9868.2005.00512.x
复制
发表时间:
2005-09-01
影响因子:
5.8
通讯作者:
Eguchi, S
Eguchi, S
中科院分区:
数学1区
文献类型:
--
作者:
Copas, J;Eguchi, S

文献摘要

被引文献

相似文献

不完整观察的数据分析问题在统计学中是非常常见的。如果我们也不确定模型的选择,那么它们就会加倍困难。我们提出了讨论此类问题的通用公式,并在模型偏差很小的假设下,对最大似然估计产生的偏差进行了近似。由于数据不完整而导致的参数估计效率损失有双重解释:假设模型正确时方差增加;当模型不正确时估计的偏差。例子包括不可忽视的缺失数据、观察研究中隐藏的混杂因素以及荟萃分析中的发表偏倚。建议在计算置信区间或检验统计量之前将方差加倍,作为解决模型中无法检测到的微小偏差的可能性的粗略方法。评估被动吸烟导致肺癌的风险的问题被用作一个激励性的例子。
Problems of the analysis of data with incomplete observations are all too familiar in statistics. They are doubly difficult if we are also uncertain about the choice of model. We propose a general formulation for the discussion of such problems and develop approximations to the resulting bias of maximum likelihood estimates on the assumption that model departures are small. Loss of efficiency in parameter estimation due to incompleteness in the data has a dual interpretation: the increase in variance when an assumed model is correct; the bias in estimation when the model is incorrect. Examples include non-ignorable missing data, hidden confounders in observational studies and publication bias in meta-analysis. Doubling variances before calculating confidence intervals or test statistics is suggested as a crude way of addressing the possibility of undetectably small departures from the model. The problem of assessing the risk of lung cancer from passive smoking is used as a motivating example.