Automatic Zig-Zag sampling in practice

Automatic Zig-Zag sampling in practice
复制标题

DOI:
10.1007/s11222-022-10142-x
复制
发表时间:
2022-06
影响因子:
2.2
通讯作者:
Alice Corbella;S. Spencer;G. Roberts
Alice Corbella;S. Spencer;G. Roberts
中科院分区:
数学2区
文献类型:
--
作者:
Alice Corbella;S. Spencer;G. Roberts

文献摘要

被引文献

相似文献

从目标分布中生成样本的新型蒙特卡罗方法,如贝叶斯分析的后验,在过去十年中得到了迅速发展。基于分段确定性马尔可夫过程(PDMPs)的非可逆连续时间过程的算法,由于其重要的特性(如超效率),正在发展成为自己的研究分支。然而,在这一领域,实践并没有跟上理论的步伐,使用PDMPs来解决应用问题的情况并不普遍。首先,这可能是由于基于pdp的采样器所面临的几个实施挑战,其次,缺乏展示应用环境中方法和实现的论文。在这里,我们使用最有前途的PDMPs之一,锯齿形采样器,作为一个原型例子来解决这两个问题。在解释了z - zag采样器的关键元素之后,暴露并解决了其实现挑战。具体地说,提供了从感兴趣的目标分布中提取样本的算法的公式。值得注意的是,该算法的唯一要求是一个封闭形式的可微函数来评估感兴趣的对数目标密度,并且,与以前的实现不同,不需要关于目标的进一步信息。通过对规范哈密顿蒙特卡罗算法的性能进行评估,证明了该算法在模拟和实际数据设置中具有竞争力。最后,我们证明了在实践中可以获得超效率特性,即以比评估所有数据的可能性更低的成本绘制一个独立样本的能力。
Novel Monte Carlo methods to generate samples from a target distribution, such as a posterior from a Bayesian analysis, have rapidly expanded in the past decade. Algorithms based on Piecewise Deterministic Markov Processes (PDMPs), non-reversible continuous-time processes, are developing into their own research branch, thanks their important properties (e.g., super-efficiency). Nevertheless, practice has not caught up with the theory in this field, and the use of PDMPs to solve applied problems is not widespread. This might be due, firstly, to several implementational challenges that PDMP-based samplers present with and, secondly, to the lack of papers that showcase the methods and implementations in applied settings. Here, we address both these issues using one of the most promising PDMPs, the Zig-Zag sampler, as an archetypal example. After an explanation of the key elements of the Zig-Zag sampler, its implementation challenges are exposed and addressed. Specifically, the formulation of an algorithm that draws samples from a target distribution of interest is provided. Notably, the only requirement of the algorithm is a closed-form differentiable function to evaluate the log-target density of interest, and, unlike previous implementations, no further information on the target is needed. The performance of the algorithm is evaluated against canonical Hamiltonian Monte Carlo, and it is proven to be competitive, in simulation and real-data settings. Lastly, we demonstrate that the super-efficiency property, i.e. the ability to draw one independent sample at a lesser cost than evaluating the likelihood of all the data, can be obtained in practice.