Ensemble Inference Methods for Models With Noisy and Expensive Likelihoods

Ensemble Inference Methods for Models With Noisy and Expensive Likelihoods
复制标题

DOI:
10.1137/21m1410853
复制
发表时间:
2022-01-01
影响因子:
2.1
通讯作者:
Wolfram, Marie-Therese
Wolfram, Marie-Therese
中科院分区:
数学3区
文献类型:
--
作者:
Dunbar, Oliver R. A.;Duncan, Andrew B.;Wolfram, Marie-Therese

文献摘要

被引文献

相似文献

数据可用性的不断提高为校准生物医学、物理和社会科学中复杂现象模型中出现的未知参数提供了机会。然而,模型的复杂性常常导致参数到数据的映射评估成本高昂,并且只能通过有噪声的近似来获得。本文关注的是利用相互作用粒子系统来解决由此产生的参数反问题。特别令人感兴趣的是,可用的正向模型评估在参数空间中会受到快速波动的影响,这种波动叠加在我们感兴趣的平滑变化的大规模参数结构之上。文中给出了一个来自气候科学的激励性示例,并且经验性地表明集合卡尔曼方法(不使用参数到数据映射的导数)表现良好。然后使用多尺度分析来分析当我们称为噪声的快速波动污染了参数到数据映射的大规模参数依赖性时,相互作用粒子系统算法的行为。从这个角度对集合卡尔曼方法和基于朗之万的方法(后者使用参数到数据映射的导数)进行了比较。结果表明,在参数到数据映射存在噪声的情况下,集合卡尔曼方法表现良好,而朗之万方法则受到不利影响。另一方面,在无噪声正向模型的设定中,朗之万方法具有正确的平衡分布,而集合卡尔曼方法除了在线性情况下,仅提供一种不受控制的近似。因此,引入了一类新的算法,即集合高斯过程采样器,它结合了集合卡尔曼方法和朗之万方法的优点,并被证明表现良好。
The increasing availability of data presents an opportunity to calibrate unknown parameters which appear in complex models of phenomena in the biomedical, physical, and social sciences. However, model complexity often leads to parameter-to-data maps which are expensive to evaluate and are only available through noisy approximations. This paper is concerned with the use of interacting particle systems for the solution of the resulting inverse problems for parameters. Of particular interest is the case where the available forward model evaluations are subject to rapid fluctuations, in parameter space, superimposed on the smoothly varying large-scale parametric structure of interest. A motivating example from climate science is presented, and ensemble Kalman methods (which do not use the derivative of the parameter-to-data map) are shown, empirically, to perform well. Multiscale analysis is then used to analyze the behavior of interacting particle system algorithms when rapid fluctuations, which we refer to as noise, pollute the large-scale parametric dependence of the parameter-to-data map. Ensemble Kalman methods and Langevin-based methods (the latter use the derivative of the parameter-to-data map) are compared in this light. The ensemble Kalman methods are shown to behave favorably in the presence of noise in the parameter-to-data map, whereas Langevin methods are adversely affected. On the other hand, Langevin methods have the correct equilibrium distribution in the setting of noise-free forward models, while ensemble Kalman methods only provide an uncontrolled approximation, except in the linear case. Therefore a new class of algorithms, ensemble Gaussian process samplers, which combine the benefits of both ensemble Kalman and Langevin methods, are introduced and shown to perform favorably.