On Optimal Selection of Summary Statistics for Approximate Bayesian Computation

On Optimal Selection of Summary Statistics for Approximate Bayesian Computation
复制标题

DOI:
10.2202/1544-6115.1576
复制
发表时间:
2010-01-01
影响因子:
0.9
通讯作者:
Balding, David J.
Balding, David J.
中科院分区:
数学4区
文献类型:
--
作者:
Nunes, Matthew A.;Balding, David J.

文献摘要

被引文献

相似文献

如何最好地总结大型和复杂的数据集是一个问题,出现在许多科学领域。我们从寻求数据摘要的角度来处理它,这些数据摘要使近似贝叶斯计算(ABC)下感兴趣的参数的后验分布的平均平方误差最小化。在ABC中,在模型下的模拟代替了似然的计算,这对于许多复杂的模型是方便的。模拟数据集和观察数据集通常使用汇总统计进行比较,通常在实践中根据研究者的直觉和该领域的既定实践进行选择。我们提出了两种算法自动选择有效的数据摘要。首先,我们激励最小化的后验近似的估计熵作为一种启发式的选择汇总统计量。其次,我们提出了一个两阶段的过程:最小熵算法被用来识别模拟数据集接近观察到的,这些都被连续视为观察数据集的ABC后验近似的均根综合平方误差最小化的汇总统计。在一个模拟研究中,我们都单独和联合推断的缩放突变和重组参数从人口样本的DNA序列。计算速度快的最小熵算法表现出适度的改进,而我们的两阶段的程序表现出实质性的和高度显着的进一步改善单变量和双变量的推断。我们发现,汇总统计量的最佳集合是高度特定于数据集的,这表明更一般地说,可能没有全局最佳选择,这就要求即使模型和推理目标不变,也要为每个数据集进行新的选择。
How best to summarize large and complex datasets is a problem that arises in many areas of science. We approach it from the point of view of seeking data summaries that minimize the average squared error of the posterior distribution for a parameter of interest under approximate Bayesian computation (ABC). In ABC, simulation under the model replaces computation of the likelihood, which is convenient for many complex models. Simulated and observed datasets are usually compared using summary statistics, typically in practice chosen on the basis of the investigator's intuition and established practice in the field. We propose two algorithms for automated choice of efficient data summaries. Firstly, we motivate minimisation of the estimated entropy of the posterior approximation as a heuristic for the selection of summary statistics. Secondly, we propose a two-stage procedure: the minimum-entropy algorithm is used to identify simulated datasets close to that observed, and these are each successively regarded as observed datasets for which the mean root integrated squared error of the ABC posterior approximation is minimized over sets of summary statistics. In a simulation study, we both singly and jointly inferred the scaled mutation and recombination parameters from a population sample of DNA sequences. The computationally-fast minimum entropy algorithm showed a modest improvement over existing methods while our two-stage procedure showed substantial and highly-significant further improvement for both univariate and bivariate inferences. We found that the optimal set of summary statistics was highly dataset specific, suggesting that more generally there may be no globally-optimal choice, which argues for a new selection for each dataset even if the model and target of inference are unchanged.