A Bootstrap Metropolis-Hastings Algorithm for Bayesian Analysis of Big Data.

A Bootstrap Metropolis-Hastings Algorithm for Bayesian Analysis of Big Data.
复制标题

DOI:
10.1080/00401706.2016.1142905
复制
发表时间:
2016
期刊:
Technometrics : a journal of statistics for the physical, chemical, and engineering sciences
影响因子:
--
通讯作者:
Song Q
Song Q
中科院分区:
其他
文献类型:
--
作者:
Liang F;Kim J;Song Q

文献摘要

参考文献

相似文献

马尔可夫链蒙特卡罗(MCMC)方法已被证明是分析复杂结构数据的一种非常强大的工具。然而,它们的计算机密集型性质通常需要大量迭代和每次迭代完整扫描整个数据集,这排除了它们用于大数据分析的可能性。在本文中,我们提出了所谓的Bootstrap Metropolis-Hastings(BMH)算法,它为如何驯服用于大数据分析的强大的MCMC方法提供了一个通用的框架,即用从多个Bootstrap样本并行计算的对数似然的蒙特卡罗平均来代替全部数据对数似然。BMH算法具有令人尴尬的并行结构,避免了在迭代中重复扫描整个数据集,因此对于大数据问题是可行的。与流行的分割合并方法相比,BMH通常更有效,因为它可以将整个数据信息渐进地集成到单个模拟运行中。BMH算法非常灵活。与Metropolis-Hastings算法一样,它可以作为开发适用于大数据问题的高级MCMC算法的基本构建块。文中以回火BMH算法为例说明了这一点,该算法可以看作是并行回火和BMH算法的结合。BMH还可以通过分别与可逆跳跃MCMC和模拟退火法相结合来进行模型选择和优化。
Markov chain Monte Carlo (MCMC) methods have proven to be a very powerful tool for analyzing data of complex structures. However, their computer-intensive nature, which typically require a large number of iterations and a complete scan of the full dataset for each iteration, precludes their use for big data analysis. In this paper, we propose the so-called bootstrap Metropolis-Hastings (BMH) algorithm, which provides a general framework for how to tame powerful MCMC methods to be used for big data analysis; that is to replace the full data log-likelihood by a Monte Carlo average of the log-likelihoods that are calculated in parallel from multiple bootstrap samples. The BMH algorithm possesses an embarrassingly parallel structure and avoids repeated scans of the full dataset in iterations, and is thus feasible for big data problems. Compared to the popular divide-and-combine method, BMH can be generally more efficient as it can asymptotically integrate the whole data information into a single simulation run. The BMH algorithm is very flexible. Like the Metropolis-Hastings algorithm, it can serve as a basic building block for developing advanced MCMC algorithms that are feasible for big data problems. This is illustrated in the paper by the tempering BMH algorithm, which can be viewed as a combination of parallel tempering and the BMH algorithm. BMH can also be used for model selection and optimization by combining with reversible jump MCMC and simulated annealing, respectively.
DOI: 10.1016/s0168-1699(99)00046-0
发表时间: 1999-12-01
影响因子: 8.3
作者:
Blackard, JA;Dean, DJ
通讯作者: Dean, DJ
DOI: 10.1162/neco_a_00466
发表时间: 2013-08-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Liang, Faming;Jin, Ick-Hoon
通讯作者: Jin, Ick-Hoon
DOI: 10.1162/08997660360675107
发表时间: 2003-08-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Liang, FM
通讯作者: Liang, FM
DOI: 10.1214/07-aos574
发表时间: 2009-04-01
影响因子: 4.5
作者:
Andrieu, Christophe;Roberts, Gareth O.
通讯作者: Roberts, Gareth O.
DOI: 10.1214/aoms/1177730150
发表时间: 1948-01-01
影响因子: --
作者:
HOEFFDING, W
通讯作者: HOEFFDING, W