Latent Normalizing Flows for Discrete Sequences

Latent Normalizing Flows for Discrete Sequences
复制标题

DOI:
--
复制
发表时间:
2019-01
期刊:
--
影响因子:
--
通讯作者:
Zachary M. Ziegler;Alexander M. Rush
Zachary M. Ziegler;Alexander M. Rush
中科院分区:
其他
文献类型:
--
作者:
Zachary M. Ziegler;Alexander M. Rush

文献摘要

被引文献

相似文献

归一化流是一类强大的连续随机变量生成模型,显示出强大的模型灵活性和非自回归生成的潜力。在对文本等离散随机变量进行建模时,也需要这些好处,但直接将标准化流应用于离散序列会带来额外的重大挑战。我们提出了一种基于 VAE 的生成模型,该模型共同学习潜在空间中基于归一化流的分布以及到观察到的离散空间的随机映射。在这种情况下,我们发现基于流的分布高度多模态至关重要。为了捕捉这一特性,我们提出了几种标准化流架构,以最大限度地提高模型灵活性。实验考虑了字符级语言建模和复调音乐生成的常见离散序列任务。我们的结果表明,基于流的自回归模型可以与可比较的自回归基线的性能相匹配,并且基于非自回归流的模型可以提高生成速度,但会降低性能。
Normalizing flows are a powerful class of generative models for continuous random variables, showing both strong model flexibility and the potential for non-autoregressive generation. These benefits are also desired when modeling discrete random variables such as text, but directly applying normalizing flows to discrete sequences poses significant additional challenges. We propose a VAE-based generative model which jointly learns a normalizing flow-based distribution in the latent space and a stochastic mapping to an observed discrete space. In this setting, we find that it is crucial for the flow-based distribution to be highly multimodal. To capture this property, we propose several normalizing flow architectures to maximize model flexibility. Experiments consider common discrete sequence tasks of character-level language modeling and polyphonic music generation. Our results indicate that an autoregressive flow-based model can match the performance of a comparable autoregressive baseline, and a non-autoregressive flow-based model can improve generation speed with a penalty to performance.