课题基金 / 基金详情

CAREER: Beyond Conditional Independence: New Model-Free Targets for High-Dimensional Inference

CAREER: Beyond Conditional Independence: New Model-Free Targets for High-Dimensional Inference
职业:超越条件独立:高维推理的新无模型目标
批准号:
2045981
负责人:
Lucas Janson
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-07-01 至 2026-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
几十年来,统计领域从简单、易于理解的模型中汲取科学见解,取得了巨大的成功。但计算技术的进步现在使研究人员能够同时测量和存储海量数据,为理解和操纵比以往任何时候都复杂得多的系统打开了大门。事实上,机器学习领域在将预测模型与这类数据集相匹配方面非常成功,但这些方法的黑箱性质使得很难使用通常的统计方法从它们中得出科学见解。事实上,不仅传统的统计方法在现代大数据环境中失败了,而且它们回答的问题甚至不再有意义,因为它们所基于的模型甚至不是近似成立的。在这项研究中,PI将首先提出新的统计方法,提出在复杂数据中有意义但仍有可解释答案的科学问题。其次,PI将找到新的统计方法来严格回答这些问题,并研究这些方法的数学和计算特性,以便尽可能有效地使用它们。最后,PI将与基因组学、微生物组和政治学领域的专家合作,使用这些方法在这些领域获得新的科学见解。在整个项目中,PI还将提供免费的统计咨询服务,以帮助更广泛的研究社区,与高中教师一起开发新课程,并为本科生和研究生提供丰富的教育和研究经验。为了了解协变量在高维回归中的重要性,随着反应进行条件独立性的假设检验越来越受欢迎。这种测试的吸引力在于,它提供了统计上严格的洞察力,无论反应如何依赖于协变量,包括当它们的关系高度非线性并包括可能是高阶的相互作用时,这种洞察都是定义明确的,在科学上是可以解释的。然而,作为非模型推理目标的条件独立性只提供了一种在某些应用中用处不大的科学洞察力。对于两种类型的数据,PI将把条件独立性扩展到新的无模型目标,这两类数据不能提供有用的推断目标,即具有高度局部依赖协变量的数据和具有成分协变量的数据。然后,PI将完全超越条件独立性,以提出新的无模型目标,而不是仅仅识别(如在变量选择中)实际上在数字上量化数据中的关系,例如协变量和响应之间的关系,或者响应的条件分布中两个协变量之间的相互作用。随着每一个新的目标,PI将开发出全新的方法,用于强大的和可证明有效的推理。拟议目标和相关方法的新颖性将提供与其他统计学领域的新联系,包括贝叶斯计算、测量理论、统计物理、实验设计、因果推理和图形模型估计。最终,这项研究旨在为研究人员提供一套新的工具,让他们超越参数目标的限制,转而利用最先进的机器学习工具,以统计原则的方式回答关于他们数据的新的和重要的问题。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The field of statistics has seen great success over many decades drawing scientific insights from simple, easy-to-understand models. But progress in computing is now allowing researchers to measure and store huge amounts of data at once, opening the door to understanding and manipulating much more complex systems than ever before. Indeed, the field of machine learning has been very successful at fitting predictive models to such data sets, but the black-box nature of these methods makes it hard to draw scientific insights from them using the usual statistical approach. In fact, not only do classical statistical methods fail in this modern big data setting, but the questions they answer no longer even make sense because the models they are based on do not hold even approximately. In this research, the PI will first come up with new statistical ways of posing scientific questions that make sense in complex data but still have interpretable answers. Second, the PI will find new statistical methods to answer those questions in a rigorous way, and study the mathematical and computational properties of these methods so that they can be used as effectively as possible. And finally, the PI will work with experts in the areas of genomics, the microbiome, and political science to use these methods to gain new scientific insights in these fields. Throughout the project, the PI will also run a free statistical consulting service to help the broader research community, develop new curricula with high school teachers, and provide enriching educational and research experiences for undergraduate and graduate students.To understand the importance of a covariate in a high-dimensional regression, it has become increasingly popular to perform a hypothesis test for conditional independence with the response. The appeal of such a test is that it provides statistically rigorous insight that is well-defined and scientifically interpretable no matter how the response depends on the covariates, including the case when their relationship is highly nonlinear and includes interactions, possibly of high order. However, conditional independence as a model-free inferential target only provides a type of scientific insight that can be of little use in some applications. The PI will extend conditional independence to new model-free targets for two types of data on which it does not provide a useful inferential target, namely, data with highly-locally-dependent covariates and data with compositional covariates. Then, the PI will move past conditional independence entirely to propose novel model-free targets that instead of just identifying (as in variable selection) actually numerically quantify relationships in the data, such as the relationship between a covariate and a response or the interaction between two covariates in the response's conditional distribution. Along with each new target, the PI will develop entirely new methods for powerful and provably valid inference. The novelty of the proposed targets and associated methods will provide for new connections with other fields of statistics including Bayesian computation, measure theory, statistical physics, experimental design, causal inference, and graphical model estimation. Ultimately, this research aims to provide a suite of new tools for researchers to move beyond the constraints of parametric targets and instead leverage state-of-the-art machine learning tools to answer novel and important questions about their data in a statistically principled way.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金