Joint Multimodal Learning with Deep Generative Models

Joint Multimodal Learning with Deep Generative Models
复制标题

DOI:
--
复制
发表时间:
2016-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Masahiro Suzuki;Kotaro Nakayama;Y. Matsuo
Masahiro Suzuki;Kotaro Nakayama;Y. Matsuo
中科院分区:
其他
文献类型:
--
作者:
Masahiro Suzuki;Kotaro Nakayama;Y. Matsuo

文献摘要

相似文献

我们研究了可以通过双向交换多种模式的深层生成模型,例如,从相应的文本中生成图像,反之亦然。最近,一些研究在深层生成模型(例如变异自动编码器(VAE))上处理多种模式。但是,这些模型通常假定模式被迫具有条件关系,即,我们只能在一个方向上生成模态。为了实现我们的目标,我们应该提取共同表示,该表示可以捕获各种方式之间的高级概念,并通过这些方式进行双向交换。如本文所述,我们提出了一个联合多模式变异自动编码器(JMVAE),其中所有模式都以关节表示独立条件。换句话说,它模拟了模式的联合分布。此外,为了能够正确地从剩余的模态中产生缺失的方式,我们开发了一种额外的方法JMVAE-KL,该方法通过减少JMVAE编码器和准备好的各自方式网络之间的差异来训练。我们的实验表明,我们提出的方法可以从多种方式获得适当的联合表示,并且比传统的VAE可以更正确地生成和重建它们。我们进一步证明JMVAE可以双向产生多种方式。
We investigate deep generative models that can exchange multiple modalities bi-directionally, e.g., generating images from corresponding texts and vice versa. Recently, some studies handle multiple modalities on deep generative models, such as variational autoencoders (VAEs). However, these models typically assume that modalities are forced to have a conditioned relation, i.e., we can only generate modalities in one direction. To achieve our objective, we should extract a joint representation that captures high-level concepts among all modalities and through which we can exchange them bi-directionally. As described herein, we propose a joint multimodal variational autoencoder (JMVAE), in which all modalities are independently conditioned on joint representation. In other words, it models a joint distribution of modalities. Furthermore, to be able to generate missing modalities from the remaining modalities properly, we develop an additional method, JMVAE-kl, that is trained by reducing the divergence between JMVAE's encoder and prepared networks of respective modalities. Our experiments show that our proposed method can obtain appropriate joint representation from multiple modalities and that it can generate and reconstruct them more properly than conventional VAEs. We further demonstrate that JMVAE can generate multiple modalities bi-directionally.