ARMS: Antithetic-REINFORCE-Multi-Sample Gradient for Binary Variables

ARMS: Antithetic-REINFORCE-Multi-Sample Gradient for Binary Variables
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
A. Dimitriev;Mingyuan Zhou
A. Dimitriev;Mingyuan Zhou
中科院分区:
其他
文献类型:
--
作者:
A. Dimitriev;Mingyuan Zhou

文献摘要

相似文献

估计二元变量的梯度是一个经常出现在各个领域的任务,例如训练离散潜在变量模型。通常使用的是基于强化的蒙特卡罗估计方法,该方法使用独立样本或负相关样本对。为了更好地利用两个以上的样本,我们提出了ARMS,一种基于反增强的多样本梯度估计器。ARMS使用联结来生成任意数量的相互对立的样本。它是无偏的,具有低方差,并推广了两种武器,我们显示为两个样本的ARMS,以及留一强化(LOORF)估计器,这是具有不相关样本的ARMS。我们在训练生成模型的几个数据集上评估了ARMS,我们的实验结果表明它优于竞争方法。我们还开发了一个版本的ARMS用于优化多样本变分界,并表明它优于VIMCO和缴械。代码是公开的。
Estimating the gradients for binary variables is a task that arises frequently in various domains, such as training discrete latent variable models. What has been commonly used is a REINFORCE based Monte Carlo estimation method that uses either independent samples or pairs of negatively correlated samples. To better utilize more than two samples, we propose ARMS, an Antithetic REINFORCE-based Multi-Sample gradient estimator. ARMS uses a copula to generate any number of mutually antithetic samples. It is unbiased, has low variance, and generalizes both DisARM, which we show to be ARMS with two samples, and the leave-one-out REINFORCE (LOORF) estimator, which is ARMS with uncorrelated samples. We evaluate ARMS on several datasets for training generative models, and our experimental results show that it outperforms competing methods. We also develop a version of ARMS for optimizing the multi-sample variational bound, and show that it outperforms both VIMCO and DisARM. The code is publicly available.