Learning Representations for Multimodal Data with Deep Belief Nets

Learning Representations for Multimodal Data with Deep Belief Nets
复制标题

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Nitish Srivastava;R. Salakhutdinov
Nitish Srivastava;R. Salakhutdinov
中科院分区:
其他
文献类型:
--
作者:
Nitish Srivastava;R. Salakhutdinov

文献摘要

被引文献

相似文献

我们提出了一种深度信念网络架构,用于学习多模态数据的联合表示。该模型在多模态输入的空间上确定概率分布,并允许从每个数据模态的条件分布中进行采样。这使得即使在某些数据模态缺失的情况下,模型也可以创建多模态表示。我们在由图像和文本组成的双模态数据上的实验结果表明,多模态DBN可以学习图像和文本输入的联合空间的良好生成模型,这对于填充丢失的数据是有用的,因此它可以用于图像标注和图像检索。我们进一步证明,使用多模态DBN发现的表示,我们的模型在区分任务上的性能可以显着优于支持向量机和LDA。
We propose a Deep Belief Network architecture for learning a joint representation of multimodal data. The model denes a probability distribution over the space of multimodal inputs and allows sampling from the conditional distributions over each data modality. This makes it possible for the model to create a multimodal representation even when some data modalities are missing. Our experimental results on bi-modal data consisting of images and text show that the Multimodal DBN can learn a good generative model of the joint space of image and text inputs that is useful for lling in missing data so it can be used both for image annotation and image retrieval. We further demonstrate that using the representation discovered by the Multimodal DBN our model can significantly outperform SVMs and LDA on discriminative tasks.