Robust and scalable Bayesian inference under model misspecification
Robust and scalable Bayesian inference under model misspecification
批准号:
2435792
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
受模型错误指定,广义贝叶斯推理和近似推理的最新进展的启发,我建议推进计算统计和概率机器学习的最新技术水平,并引入一个方法框架,算法和相关的计算和统计理论,用于在大规模设置中对可能错误指定的模型及其模型参数的宇宙进行鲁棒推理。我的研究动机是统计机器学习在医学和社会科学中的有影响力的现实应用。最近的研究集中在模型误指定一直在寻找我们如何能够强大的误指定的可能性,误指定的先验信念和数据离群值,给定一个选定的模型家庭,和一些措施的拟合或推广。这主要是一个事后的观点,试图保护和鲁棒的算法和推理对自然源的错误。该项目试图超越事后修正,并正式嵌入到我们的算法和数学思维的任何可用信息的真实数据生成过程。 例如,该项目研究了一种适用于基于模拟器的模型的鲁棒推理方法。用这种模型进行推断是具有挑战性的,因为采样是可能的,但似然函数不可用。此外,基于模拟器的模型通常描述一些复杂的物理或生物现象,因此在实践中很容易被错误指定。也就是说,他们试图提供真实世界现象的粗略近似,然而这种近似偏离真实数据生成机制的程度可能导致误导性的推断结果。最近在一系列论文中研究了模型误定,提出了贝叶斯非参数学习(NPL)框架,该框架基于不确定性应直接施加于数据生成机制而不是感兴趣的参数的想法,这通常是传统贝叶斯方法中的情况。该项目提出的方法将此框架与最大均值离散估计相结合,提供了一种适用于无似然推理的鲁棒方法。这种方法提供了一种新颖的和计算效率高的方法与理论保证。在统计推断中广泛研究的另一种类型的模型误设是自变量之一的测量误差。这个问题,也被称为,变量误差或输入不确定性问题经常出现在经济学,医学和自然科学中,在这些领域中,通常很难或不可能精确地测量真实的世界中的量。预测变量的测量误差可能导致有偏的参数估计和潜在的误导性推断结果。例如,在许多因果推理问题中,目标是估计两个随机变量之间的因果效应,因此有偏估计会导致预测变量如何影响结果变量的错误估计。这种因果效应估计问题在健康和流行病学科学中遇到,科学家们对确定性-结果关系感兴趣。该项目旨在探索这一方向,通过调整Berkson和经典的测量误差设置的NPL框架,并实证验证所提出的方法,以现实世界中的健康和营养science.Finally,该项目探讨了模型误指定的背景下分布鲁棒优化(DRO)。我们的目标是通过最坏情况分析对变量做出稳健的决策,同时考虑可能的错误设定。在这种情况下,使用NPL后验代替标准贝叶斯后验可能会导致与传统DRO方法相比不那么保守的决策,同时也会导致模型错误。
英文摘要
Inspired by recent advances in model misspecification, generalised Bayesian inference and approximate inference, I propose to advance the state of the art in computational statistics and probabilistic machine learning and introduce a methodological framework, algorithms, and associated computational and statistical theory, for performing robust inference over a universe of potentially misspecified models and their model parameters in large-scale settings. My research is motivated by impactful real-world applications of statistical machine learning in medical and social sciences. Recent research focusing on model misspecification has been looking at how can we be robust to misspecified likelihoods, misspecified prior beliefs and data outliers, given a chosen model family, and some measure of fit or generalisation. This is primarily a post-hoc perspective that attempts to protect and robustify algorithms and inference against natural sources of misspecification. The project attempts to go beyond post-hoc corrections and formally embed into our algorithms and mathematical thinking any available information about the true data generating process. For example, the project has looked at a robust inference method applicable to simulator-based models. Inference with such models is challenging as sampling is possible however the likelihood function is unavailable. Furthermore, simulator-based models often describe some complicated physical or biological phenomena and hence can be easily misspecified in practice. That is, they attempt to provide a rough approximation of a real-world phenomenon however the degree to which this approximation deviates from the true data-generating mechanism can lead to misleading inference outcomes. Model misspecification was recently examined in a series of papers suggesting the Bayesian Nonparametric Learning (NPL) framework which is based on the idea that uncertainty should be imposed directly on the data-generating mechanism rather than a parameter of interest, which is usually the case in traditional Bayesian methodology. The project's proposed method combines this framework with Maximum Mean Discrepancy estimators to provide a robust method suitable for likelihood-free inference. Such an approach provides a novel and computationally efficient method with theoretical guarantees. A different type of model misspecification, widely studied in statistical inference, is measurement error in one of the independent variables. This problem, also called, errors-in-variables or input uncertainty problem arises often in economics, medical and natural sciences in which it is often hard or impossible to measure quantities in the real world exactly. Measurement error in the predictor variable can lead to biased parameter estimates and potentially misleading inference outcomes. For example, in many causal inference problems, the aim is to estimate the causal effect between two random variables, hence a biased estimate leads to false estimates of how the predictor variable affects the outcome variable. Such causal effect estimation problems are met in health and epidemiology sciences where scientists are interested in exposure-outcome relations. This project aims to explore this direction by adapting the NPL framework for Berkson and classical measurement error settings and empirically validating the proposed methodology to real-world applications in health and nutritional sciences.Finally, the project explores model misspecification in the context of Distributionally Robust Optimisation (DRO). The goal is to make robust decisions with respect to a variable while accounting for likelihood misspecification through a worst-case analysis. The use of an NPL posterior in place of a standard Bayesian posterior in this setting can potentially result in less conservative decisions in comparison to traditional DRO methods while also accounting for model misspecification.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位: