Training Discrete Deep Generative Models via Gapped Straight-Through Estimator

Training Discrete Deep Generative Models via Gapped Straight-Through Estimator
复制标题

DOI:
10.48550/arxiv.2206.07235
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Ting-Han Fan;Ta-Chung Chi;Alexander I. Rudnicky;P. Ramadge
Ting-Han Fan;Ta-Chung Chi;Alexander I. Rudnicky;P. Ramadge
中科院分区:
其他
文献类型:
--
作者:
Ting-Han Fan;Ta-Chung Chi;Alexander I. Rudnicky;P. Ramadge

文献摘要

被引文献

相似文献

虽然深度生成模型在图像处理、自然语言处理和强化学习方面取得了成功,但由于其梯度估计过程的高方差,涉及离散随机变量的训练仍然具有挑战性。蒙特卡罗是大多数方差缩减方法中常用的解决方案。然而,这涉及到耗时的重新编译和多个函数求值。我们提出了一个间隙直通(GST)估计,以减少方差,而不会产生rescue开销。该估计器的灵感来自直通Gumbel-Softmax的基本属性。我们确定这些属性,并通过消融研究表明,它们是必不可少的。实验表明,建议的GST估计享有更好的性能相比,强基线的两个离散的深度生成建模任务,MNIST-VAE和ListOps。
While deep generative models have succeeded in image processing, natural language processing, and reinforcement learning, training that involves discrete random variables remains challenging due to the high variance of its gradient estimation process. Monte Carlo is a common solution used in most variance reduction approaches. However, this involves time-consuming resampling and multiple function evaluations. We propose a Gapped Straight-Through (GST) estimator to reduce the variance without incurring resampling overhead. This estimator is inspired by the essential properties of Straight-Through Gumbel-Softmax. We determine these properties and show via an ablation study that they are essential. Experiments demonstrate that the proposed GST estimator enjoys better performance compared to strong baselines on two discrete deep generative modeling tasks, MNIST-VAE and ListOps.