Appropriate reduction of the posterior distribution in fully Bayesian inversions

Appropriate reduction of the posterior distribution in fully Bayesian inversions
复制标题

DOI:
10.1093/gji/ggac231
复制
发表时间:
2022-07-23
影响因子:
2.8
通讯作者:
Nozue, Yohei
Nozue, Yohei
中科院分区:
地球科学2区
文献类型:
--
作者:
Sato, D.;Fukahata, Yukitoshi;Nozue, Yohei

文献摘要

被引文献

相似文献

贝叶斯反演从观测方程和先验信息生成模型参数的后验分布,两者都由超参数加权。先验也被引入到完全贝叶斯反演中的超参数中,使我们能够通过联合后验概率地评估模型参数和超参数。然而,即使在线性逆问题中,如何从联合后验中提取模型参数的有用信息也是一个未解决的问题。本文对全贝叶斯反演中联合后验数据的适当降维问题进行了理论探讨。针对联合后验的边缘化问题,我们将概率约简的方法分为三类:(1)使用无边缘化的联合后验,(2)使用模型参数的边缘后验,(3)使用超参数的边缘后验。首先,我们得出几个分析结果,这些类别的特点。一个是一套半解析表示的概率最大化估计的线性逆问题中的各个类别。对于大量的数据和模型参数,发现(1)和(2)类的模式估计量是渐近相同的。我们还证明了第(2)和第(3)类的渐近分布δ函数集中在它们的概率峰值上,这预示着模型参数的两个不同的最优估计。其次,我们进行了一个综合测试,并发现一个适当的约简是实现的类别(3),赤池的贝叶斯信息准则为代表。其他减少类别的情况下,许多模型参数,模型参数的边缘后验概率集中不再意味着中心极限定理不合适。这些结果的主要原因是,随着模型参数数量的增加,联合后验峰在欠拟合或过拟合解处急剧上升。概率空间在模型参数维度上的指数增长使得几乎零概率事件对后验均值有贡献,并且类别(1)和(2)的分布是病态的。这种病理的一种补救措施是通过在指数多重性的模型参数空间上积分联合后验来计数所有模型参数实现。因此,类别(3)的超参数的边际后验变得合适,并且即使有许多模型参数也可以符合大数定律。后验均值和ABIC估计的指数罕见性意味着普通Monte Carlo方法在总体均值和ABIC计算中的指数时间复杂性。我们还提出了一个地球物理应用程序,从空间离散的全球导航卫星系统数据估计一个连续的应变率场,证明了模型参数场的密集基函数扩展导致在朴素的完全贝叶斯方法中的过平滑估计,而详细的字段解决收敛的减少类别(3)。我们经常天真地认为,一个好的解决方案可以从有限数量的高概率样本中构造出来,但是高概率域可能是不合适的,并且指数级的许多样本对于在高维完全贝叶斯后验概率空间中生成适当的估计是必要的。
Bayesian inversion generates a posterior distribution of model parameters from an observation equation and prior information both weighted by hyperparameters. The prior is also introduced for the hyperparameters in fully Bayesian inversions and enables us to evaluate both the model parameters and hyperparameters probabilistically by the joint posterior. However, even in a linear inverse problem, it is unsolved how we should extract useful information on the model parameters from the joint posterior. This study presents a theoretical exploration into the appropriate dimensionality reduction of the joint posterior in the fully Bayesian inversion. We classify the ways of probability reduction into the following three categories focused on the marginalization of the joint posterior: (1) using the joint posterior without marginalization, (2) using the marginal posterior of the model parameters and (3) using the marginal posterior of the hyperparameters. First, we derive several analytical results that characterize these categories. One is a suite of semi-analytic representations of the probability maximization estimators for respective categories in the linear inverse problem. The mode estimators of categories (1) and (2) are found asymptotically identical for a large number of data and model parameters. We also prove the asymptotic distributions of categories (2) and (3) delta-functionally concentrate on their probability peaks, which predicts two distinct optimal estimates of the model parameters. Secondly, we conduct a synthetic test and find an appropriate reduction is realized by category (3), typified by Akaike's Bayesian information criterion. The other reduction categories are shown inappropriate for the case of many model parameters, where the probability concentration of the marginal posterior of the model parameters no longer implies the central limit theorem. The main cause of these results is that the joint posterior peaks sharply at an underfitted or overfitted solution as the number of model parameters increases. The exponential growth of the probability space in the model-parameter dimension makes almost-zero-probability events finitely contribute to the posterior mean and distributions of categories (1) and (2) be pathological. One remedy for this pathology is counting all model-parameter realizations by integrating the joint posterior over the model-parameter space of exponential multiplicity. Hence, the marginal posterior of the hyperparameters for categories (3) becomes appropriate and can conform to the law of large numbers even with numerous model parameters. The exponential rarity of the posterior mean and ABIC estimates implies the exponential time complexity of ordinary Monte Carlo methods in population mean and ABIC computations. We also present a geophysical application to estimate a continuous strain-rate field from spatially discrete global navigation satellite system data, demonstrating denser basis function expansions of the model-parameter field lead to oversmoothed estimates in naive fully Bayesian approaches, while detailed fields are resolved with convergence by the reduction of category (3). We often naively believe a good solution can be constructed from a finite number of samples with high probabilities, but the high-probability domain could be inappropriate, and exponentially many samples become necessary for generating appropriate estimates in the high-dimensional fully Bayesian posterior probability space.