An Adaptive Subsampling Approach for MCMC Inference in Large Datasets

An Adaptive Subsampling Approach for MCMC Inference in Large Datasets
复制标题

大型数据集中 MCMC 推理的自适应子采样方法

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
C. Holmes
C. Holmes
中科院分区:
--
文献类型:
--
作者:
R. Bardenet;A. Doucet;C. Holmes

文献摘要

被引文献

相似文献

马尔可夫链蒙特卡罗(MCMC)方法通常被认为计算量太大,对大数据集没有任何实际用途。在此背景下,本文描述了一种旨在扩大Metropolis-Hastings(MH)算法的方法。我们提出了一种MH的接受/拒绝步骤的近似实现,它只需要评估数据的随机子集的可能性,但保证与基于完整数据集的接受/拒绝步骤一致,并且概率高于用户指定的容差水平。这种自适应二次抽样技术是最近在[15]中发展的方法的一种替代方法,它允许我们严格地建立所产生的近似MH算法从感兴趣的目标分布的扰动版本中采样,其到该目标的总变化距离是明确控制的。我们通过几个例子探讨了该方案的优点和局限性。
Markov chain Monte Carlo (MCMC) methods are often deemed far too computationally intensive to be of any practical use for large datasets. This paper describes a methodology that aims to scale up the Metropolis-Hastings (MH) algorithm in this context. We propose an approximate implementation of the accept/reject step of MH that only requires evaluating the likelihood of a random subset of the data, yet is guaranteed to coincide with the accept/reject step based on the full dataset with a probability superior to a user-specified tolerance level. This adaptive subsampling technique is an alternative to the recent approach developed in [15], and it allows us to establish rigorously that the resulting approximate MH algorithm samples from a perturbed version of the target distribution of interest, whose total variation distance to this very target is controlled explicitly. We explore the benefits and limitations of this scheme on several examples.