A general framework to adjust for missing confounders in observational studies
A general framework to adjust for missing confounders in observational studies
批准号:
MR/M025195/1
负责人:
Marta Blangiardo
金额:
$41.16万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --
中文摘要
在观察性研究中评估风险因素/暴露X对健康结果Y的影响总是受到混淆问题的影响。队列研究是一个理想的信息来源,因为它们通常包含一组丰富的个体水平变量。然而,仅基于队列的研究可能存在选择偏差和缺乏人口代表性的问题。队列研究也可能缺乏统计能力来评估罕见的结果,以及地理或其他群体水平的差异,这限制了诸如地区水平的社会剥夺等背景因素的调查程度。常规收集的行政数据在代表性方面是一个很好的选择;然而,这些数据源通常对大量人群具有有限数量的变量,并且可能错过重要的预测因子/混杂因素,从而导致潜在的风险估计偏差。我们提出了一个综合这两种数据来源的总体框架,该框架利用了队列/调查中混杂因素的详细信息,并受益于登记处的统计能力和人口代表性。由于行政数据集包含目标人群中每个人的数据,而队列/调查通常只涵盖个人的一个子集,因此,从后一来源获得的混杂因素将被部分测量(即,登记中的某些单位将缺失),因此这种策略导致数据输入缺失。假设每一个单一的混杂因素可能被证明在计算上是不可行的,并且受到几个假设的限制,因为潜在的大量混杂因素需要考虑。我们将建立一个类似倾向得分的指数(我们称之为部分倾向得分- PPS)来总结来自队列/调查的混杂因素的值,因此我们只需要在缺失时估算一个变量。通过一个灵活的模型,该指数将被纳入流行病学分析,我们将能够提供X和Y之间因果关系的直接估计,因为所有的混杂因素都被考虑在内。我们将首先在个人层面的数据上建立我们的框架,然后将其扩展到总体水平,例如,通常用于总结流行病学风险的空间和时空变化的小区域研究(例如,用于疾病监测)或关注病因学问题(例如,揭示死亡率或发病率的环境/社会决定因素)。我们将使用贝叶斯全概率建模,它提供了一种灵活的方法,可以结合关于缺失数据机制的不同假设,并适应缺失数据的不同模式,通过现实的模拟研究,我们将评估框架的属性,并将其与其他最先进的方法进行比较。此外,将考虑两个实际案例研究。第一项研究将基于个人数据,评估在英格兰北部接触含氯水导致出生体重过低的风险。第二项研究将调查空气污染浓度和噪音暴露对英格兰和威尔士因心血管原因住院的影响,并将在小区域一级进行。通过案例研究,我们将能够揭示与仅基于人口登记数据的常用分析相比,我们提出的方法如何改变流行病学分析的结果,即暴露对健康结果的影响。这将有可能转化为卫生政策和战略的变化,以考虑到改进的、更准确的结果,并可能成为分析观察性研究的最新方法。
英文摘要
Assessing the impact of a risk factor/exposure X on a health outcome Y in observational studies is invariably subject to confounding issues. Cohort studies are an ideal source of information as they typically contain a rich set of individual level variables. Nevertheless a study based only on a cohort may suffer from problems of selection bias and lack of population representativeness. Cohort studies may also lack statistical power to assess rare outcomes, and geographical or other group-level variations which limits the extent to which contextual factors such as area level social deprivation can be investigated. Routinely collected administrative data are a good alternative in terms of representativeness; however, these data sources typically have a limited number of variables for a large population, and might miss important predictors/confounders leading to potentially biased estimation of the risks.We propose a general framework that integrating these two sources of data takes advantage of the detailed information on confounders from cohorts/surveys and benefits from the statistical power and population representativeness of the registries. This strategy entails missing data imputation as administrative datasets contain data on each individual in the target population, while cohorts/surveys typically cover only a subset of individuals, so that the confounders obtained from the latter source will be partially measured (i.e. will be missing for some of the units in the registries). Imputing each single confounder could prove computationally unfeasible and constrained to several assumptions given the potentially large number of confounders to consider.We will build a propensity score like index (which we will call Partial Propensity Score - PPS) to summarise the values of the confounders from the cohorts/surveys so we will need to impute only one variable when missing. Through a flexible model the index will be included in the epidemiological analysis and we will be able to provide a direct estimate of the causal link between X and Y as all the confounders have been taken into account.We will build our framework first on individual level data and then extend it to aggregated level, e.g. small area studies generally used to summarise spatial and spatio-temporal variations in epidemiological risks (e.g. for disease surveillance) or to focus on aetiological questions (e.g. to unveil environmental/social determinant of mortality or morbidity). We will use Bayesian full probability modelling which provides a flexible approach of incorporating different assumptions about the missing data mechanism and accommodating different patterns of missing data, and through realistic simulation studies we will evaluate the properties of the framework and compare it with other state-of-the-art methods. In addition two real case studies will be considered. The first will assess the risk of low birth weight given exposure to chlorine in water in Northern England and will be based on individual level data. The second will investigate the impact of air pollution concentration and noise exposure on hospital admissions from cardiovascular causes in England and Wales and will be at the small area level. Through the case studies we will be able to unveil how our proposed methodology changes the results of epidemiological analyses in terms of the effect of exposure on the health outcomes, compared to the commonly used analysis based on data from population registries only. This will have the potential of translating into changes in health policies and strategies to take into account the improved, more accurate results and could become the new state-of-the-art method for analysis of observational studies.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1093/biostatistics/kxx058
发表时间:
2019-01-01
期刊:
Biostatistics (Oxford, England)
影响因子:
--
作者:
[Wang Y, Pirani M, Hansell AL, Richardson S, Blangiardo M]
通讯作者:
Blangiardo M
DOI:
10.1002/bimj.201900241
发表时间:
2020-11
期刊:
Biometrical journal. Biometrische Zeitschrift
影响因子:
--
作者:
[Pirani M, Mason AJ, Hansell AL, Richardson S, Blangiardo M]
通讯作者:
Blangiardo M
DOI:
10.1002/env.2644
发表时间:
2020-07-29
期刊:
ENVIRONMETRICS
影响因子:
1.7
作者:
[Forlani, C., Bhatt, S., Blangiardo, M.]
通讯作者:
Blangiardo, M.
Using Ecological Propensity Score to Adjust for Missing Confounders in Small Area Studies
使用生态倾向评分来调整小区域研究中缺失的混杂因素
DOI:
10.48550/arxiv.1605.00814
发表时间:
2016
期刊:
影响因子:
--
作者:
[Wang Y]
通讯作者:
Wang Y
A statistical framework for the apportionment of particulate contaminants and their health effect determination
-
批准号:MR/T044713/1
-
项目类别:Research Grant
-
资助金额:$65.28万
-
财政年份:2021
-
负责人:Marta Blangiardo
-
依托单位:
海外基金