An exploration of the missing data mechanism in an Internet based smoking cessation trial

An exploration of the missing data mechanism in an Internet based smoking cessation trial
复制标题

DOI:
10.1186/1471-2288-12-157
复制
发表时间:
2012-10-15
影响因子:
4
通讯作者:
Sutton, Stephen
Sutton, Stephen
中科院分区:
医学3区
文献类型:
--
作者:
Jackson, Dan;Mason, Dan;Sutton, Stephen

文献摘要

被引文献

相似文献

背景:在戒烟试验中,缺失结果数据是很常见的。人们通常认为,所有这些缺失的数据都来自戒烟不成功的参与者(“缺失=吸烟”)。在这里,我们使用最近一项基于互联网的戒烟试验的数据来调查一组先验选择的基线变量中的哪些是遗漏的预测,以及支持和反对“遗漏=吸烟”假设的证据。方法:我们使用选择模型,该模型模拟了在给定结果和其他变量的情况下观察结果的概率。选择模型包括一个参数,对于该参数,零表示数据在随机(MAR)处丢失,而大值表示“丢失=吸烟”。我们在敏感性分析的背景下考察了基线变量预测能力的证据。我们使用关于尝试获得结果数据的次数和类型的数据来估计吸烟状态和缺失数据指示器之间的关联。结果:我们将我们的方法应用于iQuit戒烟试验数据。从敏感性分析中,我们获得了强有力的证据,表明年龄较大的参与者更有可能提供结果数据。尝试获取结果数据的次数和类型的模型证实,年龄是丢失数据的一个很好的预测因素。从这个模型中有微弱的证据表明,成功戒烟的参与者更有可能提供结果数据,但这一证据并不支持“缺失=吸烟”的假设。有缺失结果数据的参与者在试验结束时不吸烟的概率估计在0.14到0.19之间。结论:那些进行戒烟试验并希望进行假设数据为MAR的分析的人,应该收集基线变量,并将其纳入他们的模型中,这些变量被认为是缺失数据的良好预测因素,以便使这一假设更可信。然而,他们也应该考虑遗漏Not at Random(Mnar)模型的可能性,这些模型做出或允许的假设不像“遗漏=吸烟”那么极端。
Background: Missing outcome data are very common in smoking cessation trials. It is often assumed that all such missing data are from participants who have been unsuccessful in giving up smoking ("missing=smoking"). Here we use data from a recent Internet based smoking cessation trial in order to investigate which of a set of a priori chosen baseline variables are predictive of missingness, and the evidence for and against the "missing=smoking" assumption.Methods: We use a selection model, which models the probability that the outcome is observed given the outcome and other variables. The selection model includes a parameter for which zero indicates that the data are Missing at Random (MAR) and large values indicate "missing=smoking". We examine the evidence for the predictive power of baseline variables in the context of a sensitivity analysis. We use data on the number and type of attempts made to obtain outcome data in order to estimate the association between smoking status and the missing data indicator.Results: We apply our methods to the iQuit smoking cessation trial data. From the sensitivity analysis, we obtain strong evidence that older participants are more likely to provide outcome data. The model for the number and type of attempts to obtain outcome data confirms that age is a good predictor of missing data. There is weak evidence from this model that participants who have successfully given up smoking are more likely to provide outcome data but this evidence does not support the "missing=smoking" assumption. The probability that participants with missing outcome data are not smoking at the end of the trial is estimated to be between 0.14 and 0.19.Conclusions: Those conducting smoking cessation trials, and wishing to perform an analysis that assumes the data are MAR, should collect and incorporate baseline variables into their models that are thought to be good predictors of missing data in order to make this assumption more plausible. However they should also consider the possibility of Missing Not at Random (MNAR) models that make or allow for less extreme assumptions than "missing=smoking".