Latent Diffusion for Language Generation

Latent Diffusion for Language Generation
复制标题

DOI:
10.48550/arxiv.2212.09462
复制
发表时间:
2022-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Justin Lovelace;Varsha Kishore;Chao-gang Wan;Eliot Shekhtman;Kilian Q. Weinberger
Justin Lovelace;Varsha Kishore;Chao-gang Wan;Eliot Shekhtman;Kilian Q. Weinberger
中科院分区:
其他
文献类型:
--
作者:
Justin Lovelace;Varsha Kishore;Chao-gang Wan;Eliot Shekhtman;Kilian Q. Weinberger

文献摘要

被引文献

相似文献

扩散模型在建模连续数据模态(如图像、音频和视频)方面取得了巨大的成功,但在离散领域(如语言)中的应用有限。最近的尝试,以适应扩散的语言提出了扩散作为一种替代自回归语言生成。相反,我们将扩散视为一种补充方法,可以增强现有预训练语言模型的生成能力。我们证明了连续扩散模型可以在预训练的编码器-解码器模型的潜在空间中学习,使我们能够对可以用预训练的解码器解码成自然语言的连续潜在表示进行采样。我们表明,我们的潜在扩散模型在从数据分布中采样新文本方面比强自回归基线更有效,并且还可以实现可控生成。
Diffusion models have achieved great success in modeling continuous data modalities such as images, audio, and video, but have seen limited use in discrete domains such as language. Recent attempts to adapt diffusion to language have presented diffusion as an alternative to autoregressive language generation. We instead view diffusion as a complementary method that can augment the generative capabilities of existing pre-trained language models. We demonstrate that continuous diffusion models can be learned in the latent space of a pre-trained encoder-decoder model, enabling us to sample continuous latent representations that can be decoded into natural language with the pre-trained decoder. We show that our latent diffusion models are more effective at sampling novel text from data distributions than a strong autoregressive baseline and also enable controllable generation.