Missing at random assumption made more plausible: evidence from the 1958 British birth cohort

Missing at random assumption made more plausible: evidence from the 1958 British birth cohort
复制标题

DOI:
10.1016/j.jclinepi.2021.02.019
复制
发表时间:
2021-04-02
影响因子:
7.2
通讯作者:
Ploubidis, George B.
Ploubidis, George B.
中科院分区:
医学2区
文献类型:
--
作者:
Mostafa, Tarek;Narayanan, Martina;Ploubidis, George B.

文献摘要

被引文献

相似文献

目的:纵向调查中不可避免地会出现无应答现象。其结果是统计功效较低,并可能产生偏倚。我们实施了一个系统的数据驱动的方法,以确定在国家儿童发展研究(NCDS; 1958年英国出生队列)的无反应的预测。这些变量可以帮助使随机缺失假设更合理,这对缺失数据的处理有影响。研究设计和设置:我们使用11次扫描的数据确定了无应答的预测因子。(出生至55岁)(n = 17,415),采用参数回归和LASSO进行变量选择。儿童时期的社会经济背景不佳,早期心理健康状况较差,认知能力较低,成年后缺乏公民和社会参与,这些都与无反应有关。使用这些信息,沿着与其他数据从NCDS,我们能够复制的“人口分布”的教育程度和婚姻状况(来自外部数据),和原始分布的关键早期生命characteristics.Conclusion:确定的预测因子的非响应有可能提高随机假设的缺失的可验证性。它们可以直接用作原则性方法分析中的“辅助变量”,以减少由于缺失数据而导致的偏倚。(C)2021爱思唯尔公司All rights reserved.
Objective: Non-response is unavoidable in longitudinal surveys. The consequences are lower statistical power and the potential for bias. We implemented a systematic data-driven approach to identify predictors of non-response in the National Child Development Study (NCDS; 1958 British birth cohort). Such variables can help make the missing at random assumption more plausible, which has implications for the handling of missing dataStudy Design and Setting: We identified predictors of non-response using data from the 11 sweeps (birth to age 55) of the NCDS (n = 17,415), employing parametric regressions and the LASSO for variable selection.Results: Disadvantaged socio-economic background in childhood, worse mental health and lower cognitive ability in early life, and lack of civic and social participation in adulthood were consistently associated with non-response. Using this information, along with other data from NCDS, we were able to replicate the "population distribution" of educational attainment and marital status (derived from external data), and the original distributions of key early life characteristics.Conclusion: The identified predictors of non-response have the potential to improve the plausibility of the missing at random assumption. They can be straightforwardly used as "auxiliary variables" in analyses with principled methods to reduce bias due to missing data. (C) 2021 Elsevier Inc. All rights reserved.