课题基金 / 基金详情

CAREER: Statistical Inference in High Dimensions using Variational Approximations

CAREER: Statistical Inference in High Dimensions using Variational Approximations
职业:使用变分近似进行高维统计推断
批准号:
2239234
负责人:
Subhabrata Sen
金额:
$43.41万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2028-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
现代数据应用程序通常涉及包含大量观测和特征的海量数据集。为了促进真实的统计学习,迫切需要有原则的和计算效率高的统计方法。变分推理方法最近成为这种情况下的流行选择。术语“变分推理”指的是一种通用的开箱即用的策略,用于为广泛的一类问题开发统计算法。例如,这些算法被用来作为一个子例程在文本挖掘,超现实的人工文本和图像,机器翻译等,这种方法是非常有吸引力的,由于所提出的方法的计算效率,其上级的实际性能。尽管有这些优点,这些变分方法的严格保证仍然处于新生状态。该项目将为这一方法在不同环境中的有效性提供统计保障。随后,这些新的见解将被利用来为现代数据应用开发新的统计方法。拟议研究的结果将使从业者能够有信心地部署变分推理方法。此外,这些成果将为统计人员的工具包增加一套新的有原则的、计算效率高的方法。PI将在整个研究期间及以后交织他的研究和教学。特别是,PI将开发新的本科/研究生课程,重点是变分推理和指导学生(特别是那些来自代表性不足的背景),目的是向他们介绍统计和数据科学的机会。该项目将研究基于变分近似的统计推断,重点关注三个具体方面:(i)基于回归模型的朴素平均场(NMF)近似的统计推断,(ii)超越回归的NMF近似和(iii)高级平均场近似。在主题(i)下,PI将开发高维线性模型的经验贝叶斯方法,并使用NMF近似比较贝叶斯变量选择算法。主题(ii)将集中在隐马尔可夫随机场和贝叶斯神经网络的NMF近似。最后,主题(iii)将集中在某些替代的平均场近似。物理学家推测,如果数据点和特征的数量都很大且具有可比性,则NMF近似不再准确;相反,Thouless-Anderson-Palmer(TAP)近似,一种先进的平均场近似,应该有助于贝叶斯最优推理。拟议的研究将建立这一猜想的背景下,高维线性回归的比例渐近制度。所提出的方法的理论基础将依赖于不同的想法,起源于非线性大偏差(在概率和组合学中研究),自旋玻璃(在概率和统计物理学中研究)和图形模型。反过来,这些想法将与经典的统计思想(例如非参数最大似然)相结合,以开发计算效率高的方法进行高维推理。这个奖项反映了NSF的法定使命,并被认为是值得通过使用基金会的知识价值和更广泛的影响审查标准进行评估的支持。
英文摘要
Modern data applications routinely involve massive datasets comprising a multitude of observations and features. To facilitate statistical learning in real time, there is an urgent need for principled and computationally efficient statistical methodology. Variational Inference methods have recently emerged as a popular choice in this context. The term "Variational Inference" refers to a general out-of-the-box strategy to develop statistical algorithms for a wide class of problems. For example, these algorithms are used as a sub-routine in text mining, generation of hyper-realistic artificial text and images, machine translation, etc. This approach is extremely attractive due to the computational efficiency of the proposed methods, and their superior practical performance. Despite these advantages, rigorous guarantees for these variational methods are still in a nascent state. This project will develop statistical guarantees for the validity of this approach in diverse settings. Subsequently, these new insights will be exploited to develop novel statistical methodology for modern data applications. The outcome of the proposed research will allow practitioners to deploy Variational Inference methods with confidence. In addition, the outcomes will add a new set of principled, computationally efficient methods to the statistician's toolkit. The PI will interweave his research and teaching throughout the research period and beyond. In particular, the PI will develop new undergraduate/graduate courses focusing on Variational Inference and mentor students (particularly those from under-represented backgrounds) with the aim of introducing them to opportunities in statistics and data science. The proposed research and educational activities will broaden participation in STEM generally, and encourage careers in statistics and data science.This project will study statistical inference based on variational approximations focusing on three concrete thrusts: (i) Statistical inference based on the Naive Mean Field (NMF) approximation for regression models, (ii) NMF approximation beyond regression and (iii) Advanced Mean Field approximations. Under theme (i), the PI will develop empirical Bayes methodology for the high-dimensional linear model, and compare Bayesian variable selection algorithms using the NMF approximation. Theme (ii) will focus on the NMF approximation for Hidden Markov Random Fields and Bayesian Neural Networks. Finally, theme (iii) will focus on certain alternative mean-field approximations. Physicists conjecture that if the number of datapoints and features are both large and comparable, the NMF approximation is no longer accurate; instead, the Thouless-Anderson-Palmer (TAP) approximation, an advanced mean-field approximation, should facilitate Bayes optimal inference. The proposed research will establish this conjecture in the context of high-dimensional linear regression under a proportional asymptotic regime. The theoretical foundations of the proposed methodology will rest on disparate ideas originating in non-linear large deviations (studied in probability and combinatorics), spin glasses (studied in probability and statistical physics) and graphical models. In turn, these ideas will be combined with classical statistical ideas (e.g. nonparametric maximum likelihood) to develop computationally efficient methods for high-dimensional inference. This cross-pollination of ideas will generate independent follow up research directions in each domain.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金