课题基金 / 基金详情

Multiple imputation by chained equations for data that are missing not at random: methods development for randomised trials and observational studies

Multiple imputation by chained equations for data that are missing not at random: methods development for randomised trials and observational studies
通过链式方程对非随机丢失的数据进行多重插补:随机试验和观察性研究的方法开发
批准号:
MC_EX_MR/M025012/1
负责人:
Ian White
金额:
$21.25万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2016
资助国家:
英国
项目状态:
已结题
起止时间:
2016 至 --

项目摘要

项目成果

Ian White的其他基金

相似基金

相关文献

中文摘要
翻译
医学研究人员经常发现他们打算收集的一些数据无法收集:例如,由于无法联系到参与者或不愿意提供数据。这些缺失的数据给研究分析带来了问题,因为只包括提供数据的参与者可能会导致错误的结果。处理缺失数据的最常见方法假设缺失值与亚组内的观察值相似:例如,对于在时间1和2观察到体重但在时间3缺失的受试者,假设在时间3缺失的体重与在时间1和2观察到体重相似且在时间3观察到体重的受试者在时间3观察到的体重具有相同的平均值。这种方法被称为“随机缺失”,为分析提供了一个很好的起点,但不太可能完全正确:例如,在时间3时未观察到体重的参与者可能有更大的体重增加。因此,重要的是研究人员进行敏感性分析,其中对缺失数据做出不同的假设。我们的研究建议采用一种流行的方法来处理缺失数据,称为链式方程多重插补(MICE),以允许对缺失数据进行一系列假设。这种方法的思想是,使用所有变量之间的关系迭代地填充缺失值,然后多次填充,以表示缺失数据的不确定性。然而,目前的MICE方法是假设随机缺失。我们开发了一种新的方法来实现MICE方法,该方法不假设随机缺失:相反,研究人员必须通过指定子组内缺失值和观察值之间的可能平均差异来指定与随机缺失的偏差有多大。然而,我们只在理想化的环境中探索了新方法,特别是我们还没有探索它在随机试验或随时间测量结果的研究中的应用,这项工作将首先扩展统计理论,以处理随时间测量的结果,并看看该方法在随机试验中的表现如何。然后,它将扩展方法,以解决实践中遇到的各种问题:例如不同类型的变量,复杂的分析问题和非常大的数据集。这项工作将得到编写方便用户的软件的支持,以便在两个广泛使用的统计软件包中实施新方法。我们将在几个数据集的实践中实施该方法,包括雅芳父母和儿童纵向研究,我们将探索自我伤害的预测因素,以及戒烟和减肥的随机试验。缺失的自我伤害、戒烟和体重减轻数据都不太可能是随机缺失:我们将使用我们的主题专业知识来指定亚组内缺失值和观察值之间可能的平均差异范围,从而得出更合理的结论。这项工作可能会提出意想不到的理论问题,我们将解决。最后,我们相信这种方法将被广泛应用,所以我们将通过教程文章和运行课程向研究人员传播它。
英文摘要
Medical researchers often find that some data which they intended to collect could not be collected: for example, because participants could not be contacted or were unwilling to provide data. These missing data present problems in the analysis of the study, because including only participants who provided data may lead to incorrect results. The commonest way to handle missing data assumes that missing values are similar to observed values within subgroups: for example, for participants whose weight was observed at times 1 and 2 but missing at time 3, the missing weights at time 3 are assumed to have the same average as observed weights at time 3 in participants whose weights were similar at times 1 and 2 and observed at time 3. This approach is called "Missing at Random" and provides a good starting point for analysis but is unlikely to be entirely correct: for example, participants whose weight was unobserved at time 3 may have had a larger weight gain. It is therefore important for researchers to do sensitivity analyses in which different assumptions are made about the missing data. Our research proposes to adapt a popular method for handling missing data called Multiple Imputation by Chained Equations (MICE) to allow for a range of assumptions about the missing data. The idea of this approach is that missing values are filled in iteratively using the relationships between all the variables, and this is then done multiple times in order to express uncertainty about the missing data. However, at present the MICE method is done assuming Missing at Random. We have developed a new way to implement the MICE method which does not assume Missing at Random: instead, the researcher has to specify how big the departures from Missing at Random are, by specifying the likely average differences between missing values and observed values within subgroups. However, we have only explored the new method in idealised settings, and in particular we have not explored its use in randomised trials or in studies where outcomes are measured over time.The work will first extend the statistical theory to handle outcomes that are measured over time and see how well the method performs in randomised trials. It will then extend the methods to tackle a wide range of problems met in practice: for example different types of variables, complex analysis questions, and very large data sets. This work will be supported by writing user-friendly software to implement the new method in two widely used statistics packages. We will implement the method in practice in several data sets, including the Avon Longitudinal Study of Parents and Children where we will explore predictors of self-harm, and randomised trials in smoking cessation and weight loss. Missing self-harm, smoking cessation and weight loss data are all very unlikely to be Missing at Random: we will use our subject matter expertise to specify a range of likely average differences between missing values and observed values within subgroups and hence reach more defensible conclusions. This work is likely to raise unexpected theoretical issues which we will address.Finally, we believe that this method will be widely applicable, so we will disseminate it to researchers via tutorial articles and by running courses.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
Canonical Causal Diagrams to Guide the Treatment of Missing Data in Epidemiologic Studies.
在流行病学研究中指导缺失数据的治疗的规范因果图。
DOI: 10.1093/aje/kwy173
发表时间: 2018-12-01
期刊: American journal of epidemiology
影响因子: 5
作者: [Moreno-Betancur M, Lee KJ, Leacy FP, White IR, Simpson JA, Carlin JB]
通讯作者: Carlin JB
DOI: 10.1002/sim.8584
发表时间: 2020-09-30
期刊: Statistics in medicine
影响因子: 2
作者: [Tompsett D, Sutton S, Seaman SR, White IR]
通讯作者: White IR
DOI: 10.1002/jrsm.1188
发表时间: 2016-09
期刊: Research synthesis methods
影响因子: 9.8
作者: [Jackson D, Boddington P, White IR]
通讯作者: White IR
DOI: 10.1002/jrsm.1191
发表时间: 2016-09
期刊: Research synthesis methods
影响因子: 9.8
作者: [Baker R, Jackson D]
通讯作者: Jackson D
共 7 条
    Method - Design
    • 批准号:
      MC_UU_00004/09
    • 项目类别:
      Intramural
    • 资助金额:
      $275.23万
    • 财政年份:
      2021
    • 负责人:
      Ian White
    • 依托单位:
    REU Site: New approaches to engineering cells, tissues, and organs
    CAREER: Paper-based surface enhanced Raman spectroscopy (P-SERS) for biosensing using inkjet-fabricated devices
    SBIR Phase I: Extension of Multiphoton Polymerization fabrication technology to the fabrication of Retinal Image Management (RIM) elements
    • 批准号:
      0638051
    • 项目类别:
      Standard Grant
    • 资助金额:
      $0.0万
    • 财政年份:
      2007
    • 负责人:
      Ian White
    • 依托单位:
    国内基金
    海外基金
    利用Imputation和Meta分析方法深度搜寻IgA肾病新的易感基因
    • 批准号:
      81570599
    • 项目类别:
      面上项目
    • 资助金额:
      57.0万元
    • 批准年份:
      2015
    • 负责人:
      李明
    • 依托单位:
    数据缺失时高维数据降维分析的方法、理论与应用
    Imputation法及其在MHC区域易感基因搜寻中的应用
    • 批准号:
      31000528
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      19.0万元
    • 批准年份:
      2010
    • 负责人:
      左先波
    • 依托单位: