Show, Attend and Tell: Neural Image Caption Generation with Visual Attention

Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
复制标题

DOI:
--
复制
发表时间:
2015-02
期刊:
--
影响因子:
--
通讯作者:
Ke Xu;Jimmy Ba;Ryan Kiros;Kyunghyun Cho;Aaron C. Courville;R. Salakhutdinov;R. Zemel;Yoshua Bengio-Yoshua-Ben
Ke Xu;Jimmy Ba;Ryan Kiros;Kyunghyun Cho;Aaron C. Courville;R. Salakhutdinov;R. Zemel;Yoshua Bengio-Yoshua-Ben
中科院分区:
其他
文献类型:
--
作者:
Ke Xu;Jimmy Ba;Ryan Kiros;Kyunghyun Cho;Aaron C. Courville;R. Salakhutdinov;R. Zemel;Yoshua Bengio-Yoshua-Ben

文献摘要

被引文献

相似文献

受机器翻译和对象检测领域近期工作的启发,我们引入了一种基于注意力的模型,该模型可以自动学习描述图像内容。我们描述了如何使用标准的反向传播技术和随机算法,通过最大化变分下限,以确定性的方式训练这个模型。我们还通过可视化展示了模型如何能够自动学习将目光固定在突出对象上,同时在输出序列中生成相应的单词。我们在三个基准数据集上验证了注意力的使用:Flickr9k,Flickr30k和MS COCO。
Inspired by recent work in machine translation and object detection, we introduce an attention based model that automatically learns to describe the content of images. We describe how we can train this model in a deterministic manner using standard backpropagation techniques and stochastically by maximizing a variational lower bound. We also show through visualization how the model is able to automatically learn to fix its gaze on salient objects while generating the corresponding words in the output sequence. We validate the use of attention with state-of-the-art performance on three benchmark datasets: Flickr9k, Flickr30k and MS COCO.