Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models
复制标题

DOI:
--
复制
发表时间:
2014-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Ryan Kiros;R. Salakhutdinov;R. Zemel
Ryan Kiros;R. Salakhutdinov;R. Zemel
中科院分区:
其他
文献类型:
--
作者:
Ryan Kiros;R. Salakhutdinov;R. Zemel

文献摘要

被引文献

相似文献

受多模态学习和机器翻译最新进展的启发,我们引入了一个编码器-解码器管道,它学习(a):具有图像和文本的多模态联合嵌入空间和(B):一种用于从我们的空间解码分布式表示的新型语言模型。我们的管道有效地将联合图像-文本嵌入模型与多模态神经语言模型统一起来。我们介绍了结构-内容神经语言模型,该模型将句子的结构与其内容分离,以编码器产生的表示为条件。编码器允许对图像和句子进行排序,而解码器可以从头开始生成新颖的描述。使用LSTM对句子进行编码,我们在不使用对象检测的情况下,在Flickr8K和Flickr30K上匹配了最先进的性能。我们还在使用19层牛津卷积网络时设置了新的最佳结果。此外,我们表明,使用线性编码器,学习的嵌入空间捕获多模态的向量空间算术,例如 * 蓝色汽车的图像 *-"蓝色"+"红色"是红色汽车的图像附近。为800幅图像生成的示例字幕可供比较。
Inspired by recent advances in multimodal learning and machine translation, we introduce an encoder-decoder pipeline that learns (a): a multimodal joint embedding space with images and text and (b): a novel language model for decoding distributed representations from our space. Our pipeline effectively unifies joint image-text embedding models with multimodal neural language models. We introduce the structure-content neural language model that disentangles the structure of a sentence to its content, conditioned on representations produced by the encoder. The encoder allows one to rank images and sentences while the decoder can generate novel descriptions from scratch. Using LSTM to encode sentences, we match the state-of-the-art performance on Flickr8K and Flickr30K without using object detections. We also set new best results when using the 19-layer Oxford convolutional network. Furthermore we show that with linear encoders, the learned embedding space captures multimodal regularities in terms of vector space arithmetic e.g. *image of a blue car* - "blue" + "red" is near images of red cars. Sample captions generated for 800 images are made available for comparison.