Statistical Procedures and Performance Measures for Simulator-Based Frequentist Inference
Statistical Procedures and Performance Measures for Simulator-Based Frequentist Inference
批准号:
2053804
负责人:
Ann Lee
金额:
$42.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-07-01 至 2024-06-30
中文摘要
物理、工程和生物科学的许多领域都广泛使用计算机模拟器来模拟复杂的系统。虽然这些模拟器可能能够生成真实的合成数据,但它们通常不适合推断与观察到的现实世界现象相关的潜在科学机制的反问题。因此,科学界最近的一个趋势是将近似模型拟合到高保真模拟器上,然后使用这些近似模型进行科学推断。不可避免的是,任何下游分析都将依赖于近似模型、收集的数据以及模拟设计的可信度。该项目将通过为基于模拟器的科学推理和不确定性量化提供改进的程序和新的性能度量,推进理解复杂物理系统的统计方法。我们的工作将促进以数据为中心的合作的发展,并在广泛的科学领域培养学生,进一步扩大我们在高能物理、大气科学、气候学和天文学方面正在进行的跨学科研究工作。参数估计、置信集和假设检验是统计推断的标志。执行这些任务的传统方法有时不能应用于物理科学中的问题,因为(i)数据设置复杂,(ii)唯一有意义的模型是高保真正演模拟器。例如,在高能物理学中,寻找新的相互作用和粒子需要进行假设测试,包括模拟高维碰撞事件及其与粒子探测器的相互作用;在宇宙学中,科学家经常使用大n体模拟来了解宇宙是如何形成和演化的;在大气科学中,基于卫星观测推断陆地-空气碳通量依赖于复杂的大气输送模型。一个关键问题是,当连接底层参数和可观测数据的似然函数难以处理,但可以从隐式似然模型向前模拟可观测数据时,是否仍然可以构建具有适当频率覆盖和高功率的假设检验和置信集。一个相关的问题是如何校准和评估适合高保真仿真的代理模型的性能。本项目旨在通过以下目标设计将经典统计学与现代机器学习(例如,深度生成模型,神经网络分类器和凸优化)统一起来的统计程序:(1)在基于模拟器的推理设置中构建具有有限样本有效性的统计测试和频率置信集的可扩展工具和理论;(2)统计上严格的验证方法,可以在特征和参数空间上以统计置信度量化和诊断高维数据的拟合模型的质量;(3)顺序测试策略,使我们能够确定如何最好地模拟数据,以改进目标1中的测试和置信度集。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Many areas of the physical, engineering and biological sciences make extensive use of computer simulators to model complex systems. Whereas these simulators may be able to generate realistic synthetic data, they are often poorly suited for the inverse problem of inferring the underlying scientific mechanisms associated with observed real-world phenomena. Hence, a recent trend in the sciences has been to fit approximate models to high-fidelity simulators, and then use these approximate models for scientific inference. Inevitably, any downstream analysis will depend on the trustworthiness of the approximate model, the data collected, as well as the design of the simulations. This project will advance statistical methods for understanding complex physical systems by providing improved procedures and new performance measures for simulator-based scientific inference and uncertainty quantification. Our work will stimulate the development of data-focused collaborations, and training of students across a wide range of scientific areas, further expanding upon our ongoing interdisciplinary research efforts in high-energy physics, atmospheric science, climatology, and astronomy.Parameter estimation, confidence sets, and hypothesis testing are the hallmarks of statistical inference. Traditional methods to perform such tasks can sometimes not be applied to problems in the physical sciences because of (i) complex data settings, and (ii) the only meaningful model existing as a high-fidelity forward simulator. For example, in high-energy physics, searches of new interactions and particles require hypothesis tests involving simulations of high-dimensional collision events and their interactions with particle detectors; in cosmology, scientists regularly use large N-body simulations to understand how the Universe formed and evolved; and in atmospheric science, inferring land-air carbon fluxes based on satellite observations relies on complex atmospheric transport models. A key question is whether one can still construct hypothesis tests and confidence sets with proper frequentist coverage and high power when the likelihood function, which connects underlying parameters with observable data, is intractable but one can forward-simulate observable data from an implicit likelihood model. A related question is how to calibrate and assess the performance of surrogate models fit to high-fidelity simulations. This project works toward designing statistical procedures that unify classical statistics with modern machine learning (e.g., deep generative models, neural network classifiers and convex optimization) via the following aims: (1) Scalable tools and theory for constructing statistical tests and frequentist confidence sets with finite-sample validity in a simulator-based inference setting; (2) Statistically rigorous validation methods, which can quantify and diagnose the quality of fitted models of high-dimensional data with statistical confidence across both feature and parameter space; and (3) Sequential testing strategies that allow us to identify how to best simulate data to improve tests and confidence sets in Aim 1.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(12)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2305.00070
发表时间:
2023-04
期刊:
影响因子:
--
作者:
[Chirag Gupta;Aaditya Ramdas]
通讯作者:
Chirag Gupta;Aaditya Ramdas
DOI:
--
发表时间:
2021-07
期刊:
影响因子:
--
作者:
[Chirag Gupta;Aaditya Ramdas]
通讯作者:
Chirag Gupta;Aaditya Ramdas
Statistical constraints on climate model parameters using a scalable cloud-based inference framework
使用可扩展的基于云的推理框架对气候模型参数进行统计约束
DOI:
10.1017/eds.2023.12
发表时间:
2023
期刊:
Environmental Data Science
影响因子:
--
作者:
[Carzon, James, Abreu, Bruno, Regayre, Leighton, Carslaw, Kenneth, Deaconu, Lucia, Stier, Philip, Gordon, Hamish, Kuusela, Mikael]
通讯作者:
Kuusela, Mikael
DOI:
--
发表时间:
2022-05
期刊:
影响因子:
--
作者:
[Luca Masserano;T. Dorigo;Rafael Izbicki;Mikael Kuusela;Ann B. Lee]
通讯作者:
Luca Masserano;T. Dorigo;Rafael Izbicki;Mikael Kuusela;Ann B. Lee
DOI:
--
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
作者:
[Ziyu Xu;Ruodu Wang;Aaditya Ramdas]
通讯作者:
Ziyu Xu;Ruodu Wang;Aaditya Ramdas
共 12 条
Complexity to Clarity: Nonparametric Procedures that Exploit Structured Data and Models
-
批准号:1521786
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2015
-
负责人:Ann Lee
-
依托单位:
MSPA - AST: Sparse Representation and Efficient Inference for Astronomical Spectra
-
批准号:0707059
-
项目类别:Standard Grant
-
资助金额:$24.0万
-
财政年份:2007
-
负责人:Ann Lee
-
依托单位:
International Research Fellow Awards Program: Biomechanical Regulation of Cardiovascular Collagen
-
批准号:9600380
-
项目类别:Fellowship Award
-
资助金额:$2.42万
-
财政年份:1996
-
负责人:Ann Lee
-
依托单位:
海外基金