Instrumental Variables Estimation, Selection of Instruments and Two-Sample Mendelian Randomisation
Instrumental Variables Estimation, Selection of Instruments and Two-Sample Mendelian Randomisation
批准号:
2740743
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
了解一个变量如何导致另一个变量是所有科学的广泛兴趣和重要性。一个例子是肥胖对心脏病的因果影响。许多人认为进行随机实验是唯一的解决方案,但由于伦理,财务或实际考虑,这并不总是可行的。工具变量估计是一种从观测数据中识别和估计因果效应的方法。第三个变量是工具,它影响的是暴露,而不是结果,除非是通过它对暴露的影响。它能够消除由影响暴露和结果的未观察变量引起的偏倚,因此具有广泛的价值。这种情况称为未观察到的混杂。在观察性遗传研究中,孟德尔随机化(MR)采用遗传变异作为工具变量,以发现可改变的健康暴露对疾病的因果影响。双样本MR使用两个独立的数据集通过汇总数据进行统计推断,通过不需要访问个人水平的信息来维护隐私。本研究旨在开发新的统计方法来改进因果推断中的工具变量估计,以解决实际应用中的有偏估计。最流行的MR估计是逆方差加权(IVW)估计。然而,由于许多弱工具的问题,IVW可能会严重偏倚。引入预选步骤已成为实践中的标准。虽然这种选择消除了许多弱工具的偏见,但它带来了另一种偏见,赢家的诅咒偏见,因为幸存的工具是赢家。找到第三个独立的样本解决了这种偏差,但依靠这样的样本往往是不可行的。因此,寻找一个更好的估计器已经引起了研究者的关注。最近的建议包括去偏IVW(dIVW)估计和轮廓分数(PS)估计。dIVW用作IVW估计器的增强迭代。在评估估计量时,两个最重要的性质是最重要的:一致性和渐近正态性,特别是在有大量工具的情况下。尽管这三种估计都具有这些特性,但它们依赖于不同的假设集来有效地发挥作用。本项目的中心目标是深入审查这些假设,以便加以完善,并为选择最合适的估计数制定标准。更具体地说,该项目需要比较分析的条件下,dIVW和PS的估计优于。据设想,该项目将产生理论见解和经验验证的混合。该项目还涉及有效工具的选择,由两个排除条件定义-(1)与结果不直接相关,(2)与影响暴露和结果的未观察到的变量不相关。现有的置信区间方法只能从假设满足条件(2)的工具中选择满足条件(1)的工具。该项目计划开发一种方法来选择有效的工具,而不必假设条件(1)或(2)。一个研究方向专注于双/去偏机器学习。许多研究主要集中在线性模型规格。我将允许模型中的一般功能,并训练机器学习这些功能。此外,我将使用正交矩来消除估计量的偏差,并使用交叉拟合来消除过拟合偏差,并改善推理。还有更多可能的方向,例如变化(异方差)方差或用于在工具选择中选择阈值的算法。解决这些问题,结合必要的软件开发,具有重要的相关性,并准备在应用因果推理领域内受益于广大的用户群。该项目属于EPSRC数学科学领域的福尔斯。
英文摘要
Learning about how one variable causes the other is of broad interest and importance across all sciences. An example is the causal effect of obesity on heart disease. Many consider conducting randomised experiments the only solution, but that is not always feasible due to ethical, financial or practical considerations.Instrumental Variables estimation is a method to identify and estimate causal effects from observational data. An instrument, a third variable, affects the exposure, not the outcome, except through its impact on the exposure. It is able to remove bias induced by an unobserved variable affecting both the exposure and outcome, thus being widely valuable. This scenario is called unobserved confounding. In observational genetic studies, Mendelian randomisation (MR) employs genetic variants as instrumental variables to discover the causal effects of modifiable health exposures on disease. Two-sample MR uses two independent data sets for statistical inference through summary data, maintaining privacy by not requiring access to individual-level information.The research aims to develop novel statistical methods for improving instrumental variables estimation in causal inference to address biased estimates in practical applications.The most popular MR estimator is the inverse variance weighted (IVW) estimator. However, IVW can be heavily biased due to the many weak instruments problem. Introducing a pre-selection step has become the standard in practice. Although such selection removes much of the weak instruments bias, it brings another bias, the winner's curse bias, because the surviving instruments are the winners. Finding a third independent sample solves such bias, but relying on such a sample is often infeasible. Therefore, finding a better estimator has gathered attention from researchers. Recent proposals include the debiased IVW (dIVW) estimator and the Profile Score (PS) estimator. dIVW serves as an enhanced iteration of the IVW estimator. When evaluating estimators, two paramount properties come to the forefront: consistency and asymptotic normality, particularly in scenarios with a large number of instruments. Although all three estimators share these properties, they rely on disparate sets of assumptions to function effectively. The central objective of this project is to undertake an in-depth examination of these assumptions, with the aim of refining them and establishing criteria for selecting the most suitable estimator. More specifically, the project entails a comparative analysis of the conditions under which the dIVW and PS estimators excel. It is envisioned that the project will yield a blend of theoretical insights and empirical validation. The project also concerns the selection of valid instruments, defined by two exclusion conditions - (1) not relate directly to the outcome and (2) not relate to unobserved variables that affect both the exposure and outcome. The existing Confidence Interval method is only able to select instruments that satisfy condition (1) from instruments assumed to satisfy condition (2). The project plans to develop a methodology to select valid instruments without having to assume conditions (1) or (2). A research direction focuses on double/debiased machine learning. Many studies primarily focus on linear model specifications. I will allow for general functions in the model and train the machine to learn about these. In addition, I shall use orthogonal moments to eliminate the bias of estimators and cross-fitting to remove over-fitting biases, and improve inference. There are many more possible directions, such as varying (heteroscedastic) variances or algorithms for choosing a threshold in instrument selection. Addressing these inquiries, in conjunction with the requisite software development, holds significant relevance and is poised to benefit a vast user base within the realm of applied causal inference. This project falls within the EPSRC mathematical science area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Errors-In-Variables模型的贝叶斯估计理论研究
-
批准号:41774009
-
项目类别:面上项目
-
资助金额:69.0万元
-
批准年份:2017
-
负责人:方兴
-
依托单位: