课题基金 / 基金详情

CAREER: Statistical Inference in Observational Studies -- Theory, Methods, and Beyond

CAREER: Statistical Inference in Observational Studies -- Theory, Methods, and Beyond
职业:观察研究中的统计推断——理论、方法及其他
批准号:
2338760
负责人:
Rajarshi Mukherjee
金额:
$45.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-07-01 至 2029-06-30

项目摘要

项目成果

Rajarshi Mukherjee的其他基金

相似基金

相关文献

中文摘要
翻译
因果推理是指从经验观察中破译实体之间因果关系的一种系统方法--一种支撑过去、现在和未来科学和社会发展的认知框架。对于设计因果推断的统计方法,金标准与随机临床试验有关,在随机临床试验中,研究人员基于纯粹的随机机制将治疗/暴露分配给研究对象。随机分配否定了由于被称为混杂因素的未知共同因素而观察到的治疗/暴露与结果之间的系统偏差。然而,随机临床试验往往是不可行的、昂贵的,而且在伦理上具有挑战性。相比之下,现代技术进步为收集各种可能性的海量数据铺平了道路,这些可能性包括健康结果、环境污染、医疗索赔、教育政策干预和基因突变等。由于解释这类数据中的混杂因素是进行有效因果推理的基本方面,现代因果推理研究的主要焦点之一是设计程序来解释复杂的混杂结构,而不预先指定不现实的统计模型。尽管在这一论述中存在着大量的方法,但在根据任意混杂因素进行调整的同时,推断暴露对结果的因果影响的最佳统计方法的完整图景仍然很大程度上是开放的。此外,在因果推断领域,有几种普遍使用的方法需要严格的理论论证和随后的修改,以便进行可重复的统计研究。该项目的动机是解决这些差距,并将分为两个相互关联的广泛主题。在第一部分,这个项目提供了第一个严格的理论镜头,以寻找疾病的因果变异的大规模遗传研究中最流行的混杂调节方法。这将反过来提出关于最佳统计因果推断程序的更深层次的问题,该程序将在项目的第二部分中进行探索。由于该项目旨在将统计学方法、概率论、计算机科学和机器学习的思想联系起来,它将提供独特的学习机会来设计新的课程和课程。因此,该项目将通过课程开发、为本科生和研究生,特别是来自代表性不足群体的学生提供研究指导,以及暑期课程,将研究与教育相结合。该项目将侧重于两个广泛且相互关联的主题,其动机是利用现代观测数据进行统计和因果推理。该项目的第一部分涉及提供基因组范围关联研究中最流行的基于主成分的种群分层调整方法的第一个详细的理论图景。该项目的这一部分还旨在提供新的方法,以纠正现有方法中现有的和以前未知的可能存在的偏见,并为从业人员在方法和研究设计之间进行选择提供指导。通过认识到大规模遗传数据分析的基本原理是识别疾病表型的因果遗传决定因素,该项目的第二部分开发了在稀疏条件下的高维模型和光滑条件下的非参数模型中因果效应的最佳统计推断的第一个完整图景。此外,该项目的这一部分回应了调整用于估计滋扰函数的学习算法的基本问题,例如结果回归和因果关系估计的倾向分数,以优化因果关系估计的下游均方误差,而不是与这些回归函数相关的预测误差。整个研究将结合高维统计推断、随机矩阵理论、高阶半参数方法和信息论的思想。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Causal inference refers to a systematic way of deciphering causal relationships between entities from empirical observations – an epistemic framework that underlies past, present, and future scientific and social development. For designing statistical methods for causal inference, the gold standard pertains to randomized clinical trials where the researcher assigns treatment/exposure to subjects under study based on pure chance mechanisms. The random assignment negates systematic bias between the observed relationship between the treatment/exposure and outcome due to unknown common factors referred to as confounders. However, randomized clinical trials are often infeasible, expensive, and ethically challenging. In contrast, modern technological advancement has paved the way for the collection of massive amounts of data across a spectrum of possibilities such as health outcomes, environmental pollution, medical claims, educational policy interventions, and genetic mutations among many others. Since accounting for confounders in such data is the fundamental aspect of conducting valid causal inference, one of the major foci of modern causal inference research have been to design procedures to account for complex confounding structures without pre-specifying unrealistic statistical models. Despite the existence of a large canvas of methods in this discourse, the complete picture of the best statistical methods for inferring the causal effect of an exposure on an outcome while adjusting for arbitrary confounders remains largely open. Moreover, there are several popularly used methods that require rigorous theoretical justification and subsequent modification for reproducible statistical research in the domain of causal inference. This project is motivated by addressing these gaps and will be divided into two broad interconnected themes. In the first part, this project provides the first rigorous theoretical lens to the most popular method of confounder adjustment in large-scale genetic studies to find causal variants of diseases. This will in turn bring forth deeper questions about optimal statistical causal inference procedures that will be explored in the second part of the project. Since the project is designed to connect ideas from across statistical methods, probability theory, computer science, and machine learning, it will provide unique learning opportunities to design new courses and discourses. The project will therefore integrate research with education through course development, research mentoring for undergraduate and graduate students, especially those from underrepresented groups, and summer programs.This project will focus on two broad and interrelated themes tied together by the motivation of conducting statistical and causal inference with modern observational data. The first part of the project involves providing the first detailed theoretical picture of the most popular principal component-based method of population stratification adjustment in genome-wide association studies. This part of the project also aims to provide new methodologies to correct for existing and previously unknown possible biases in the existing methodology as well as guidelines for practitioners for choosing between methods and design of studies. By recognizing the fundamental tenet of large-scale genetic data analysis as the identification of causal genetic determinants of disease phenotypes, the second part of the project develops the first complete picture of optimal statistical inference of causal effects in both high-dimensional under sparsity and nonparametric models under smoothness conditions. Moreover, this part of the project responds to the fundamental question of tuning learning algorithms for estimating nuisance functions, such as outcome regression and propensity score for causal effect estimation, to optimize the downstream mean-squared error of causal effect estimates instead of prediction errors associated with these regression functions. The overall research will connect ideas from high-dimensional statistical inference, random matrix theory, higher-order semiparametric methods, and information theory.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Causal Inference and Machine Learning Methods
  • 批准号:
    1941419
  • 项目类别:
    Standard Grant
  • 资助金额:
    $12.07万
  • 财政年份:
    2020
  • 负责人:
    Rajarshi Mukherjee
  • 依托单位:
海外基金