CAREER: Statistical Inference in Observational Studies -- Theory, Methods, and Beyond
CAREER: Statistical Inference in Observational Studies -- Theory, Methods, and Beyond
批准号:
2338760
负责人:
Rajarshi Mukherjee
金额:
$45.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-07-01 至 2029-06-30
中文摘要
因果推理是一种从经验观察中解读实体之间因果关系的系统方法-一种支撑过去,现在和未来科学和社会发展的认识框架。为了设计因果推断的统计方法,金标准适用于随机临床试验,其中研究人员基于纯粹的机会机制为研究对象分配治疗/暴露。由于未知的共同因素(称为混杂因素),随机分配否定了治疗/暴露与结局之间观察到的关系之间的系统偏倚。然而,随机临床试验往往是不可行的,昂贵的,道德上的挑战。相比之下,现代技术进步为收集大量数据铺平了道路,这些数据涉及一系列可能性,如健康结果、环境污染、医疗索赔、教育政策干预和基因突变等。由于在这些数据中的混杂因素的会计进行有效的因果推理的基本方面,现代因果推理研究的主要焦点之一是设计程序,以占复杂的混杂结构,而不预先指定不切实际的统计模型。尽管本文中存在大量方法,但用于推断暴露对结果的因果影响同时调整任意混杂因素的最佳统计方法的完整情况在很大程度上仍然是开放的。此外,有几种常用的方法,需要严格的理论论证和随后的修改,在因果推理领域的可重复的统计研究。该项目的动机是解决这些差距,并将分为两个广泛的相互关联的主题。在第一部分,该项目提供了第一个严格的理论透镜,以最流行的方法,混杂调整在大规模的遗传研究,以找到疾病的因果变异。这反过来又会带来更深层次的问题,最佳的统计因果推理程序,将探讨在该项目的第二部分。由于该项目旨在连接来自统计方法,概率论,计算机科学和机器学习的想法,它将提供独特的学习机会来设计新的课程和话语。 因此,该项目将通过课程开发,对本科生和研究生(特别是来自代表性不足群体的学生)的研究指导以及暑期课程,将研究与教育结合起来。该项目将侧重于两个广泛而相互关联的主题,这些主题通过利用现代观测数据进行统计和因果推理的动机联系在一起。该项目的第一部分涉及提供全基因组关联研究中最流行的基于主成分的人口分层调整方法的第一个详细的理论图片。该项目的这一部分还旨在提供新的方法,以纠正现有方法中现有的和以前未知的可能偏差,并为从业人员在方法和研究设计之间进行选择提供指导。通过认识到大规模遗传数据分析的基本原则是识别疾病表型的因果遗传决定因素,该项目的第二部分开发了第一个完整的图片,在高维稀疏和光滑条件下的非参数模型中因果效应的最佳统计推断。此外,该项目的这一部分回应了调整学习算法的基本问题,用于估计滋扰函数,例如因果效应估计的结果回归和倾向得分,以优化因果效应估计的下游均方误差,而不是与这些回归函数相关的预测误差。该奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Causal inference refers to a systematic way of deciphering causal relationships between entities from empirical observations – an epistemic framework that underlies past, present, and future scientific and social development. For designing statistical methods for causal inference, the gold standard pertains to randomized clinical trials where the researcher assigns treatment/exposure to subjects under study based on pure chance mechanisms. The random assignment negates systematic bias between the observed relationship between the treatment/exposure and outcome due to unknown common factors referred to as confounders. However, randomized clinical trials are often infeasible, expensive, and ethically challenging. In contrast, modern technological advancement has paved the way for the collection of massive amounts of data across a spectrum of possibilities such as health outcomes, environmental pollution, medical claims, educational policy interventions, and genetic mutations among many others. Since accounting for confounders in such data is the fundamental aspect of conducting valid causal inference, one of the major foci of modern causal inference research have been to design procedures to account for complex confounding structures without pre-specifying unrealistic statistical models. Despite the existence of a large canvas of methods in this discourse, the complete picture of the best statistical methods for inferring the causal effect of an exposure on an outcome while adjusting for arbitrary confounders remains largely open. Moreover, there are several popularly used methods that require rigorous theoretical justification and subsequent modification for reproducible statistical research in the domain of causal inference. This project is motivated by addressing these gaps and will be divided into two broad interconnected themes. In the first part, this project provides the first rigorous theoretical lens to the most popular method of confounder adjustment in large-scale genetic studies to find causal variants of diseases. This will in turn bring forth deeper questions about optimal statistical causal inference procedures that will be explored in the second part of the project. Since the project is designed to connect ideas from across statistical methods, probability theory, computer science, and machine learning, it will provide unique learning opportunities to design new courses and discourses. The project will therefore integrate research with education through course development, research mentoring for undergraduate and graduate students, especially those from underrepresented groups, and summer programs.This project will focus on two broad and interrelated themes tied together by the motivation of conducting statistical and causal inference with modern observational data. The first part of the project involves providing the first detailed theoretical picture of the most popular principal component-based method of population stratification adjustment in genome-wide association studies. This part of the project also aims to provide new methodologies to correct for existing and previously unknown possible biases in the existing methodology as well as guidelines for practitioners for choosing between methods and design of studies. By recognizing the fundamental tenet of large-scale genetic data analysis as the identification of causal genetic determinants of disease phenotypes, the second part of the project develops the first complete picture of optimal statistical inference of causal effects in both high-dimensional under sparsity and nonparametric models under smoothness conditions. Moreover, this part of the project responds to the fundamental question of tuning learning algorithms for estimating nuisance functions, such as outcome regression and propensity score for causal effect estimation, to optimize the downstream mean-squared error of causal effect estimates instead of prediction errors associated with these regression functions. The overall research will connect ideas from high-dimensional statistical inference, random matrix theory, higher-order semiparametric methods, and information theory.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Causal Inference and Machine Learning Methods
-
批准号:1941419
-
项目类别:Standard Grant
-
资助金额:$12.07万
-
财政年份:2020
-
负责人:Rajarshi Mukherjee
-
依托单位:
海外基金