Image Captioning with Deep Bidirectional LSTMs

Image Captioning with Deep Bidirectional LSTMs
复制标题

DOI:
10.1145/2964284.2964299
复制
发表时间:
2016-04
期刊:
Proceedings of the 24th ACM international conference on Multimedia
影响因子:
--
通讯作者:
Cheng Wang;Haojin Yang;Christian Bartz;C. Meinel
Cheng Wang;Haojin Yang;Christian Bartz;C. Meinel
中科院分区:
其他
文献类型:
--
作者:
Cheng Wang;Haojin Yang;Christian Bartz;C. Meinel

文献摘要

被引文献

相似文献

这项工作提出了一种用于图像字幕的端到端可训练深度双向 LSTM(长短期记忆)模型。我们的模型建立在深度卷积神经网络 (CNN) 和两个独立的 LSTM 网络的基础上。它能够通过利用高级语义空间中的历史和未来上下文信息来学习长期的视觉语言交互。提出了两种新颖的深度双向变体模型,其中我们以不同的方式增加非线性转换的深度,以学习分层视觉语言嵌入。提出了多裁剪、多尺度和垂直镜像等数据增强技术来防止训练深度模型时的过度拟合。我们可视化双向 LSTM 内部状态随时间的演变,并定性分析我们的模型如何将图像“翻译”为句子。我们提出的模型使用三个基准数据集:Flickr8K、Flickr30K 和 MSCOCO 数据集在标题生成和图像句子检索任务上进行评估。我们证明,即使没有集成额外的机制(例如对象检测、注意力模型等),双向 LSTM 模型在字幕生成方面也能实现与最先进的结果相比具有高度竞争力的性能,并且在检索任务上显着优于最新的方法
This work presents an end-to-end trainable deep bidirectional LSTM (Long-Short Term Memory) model for image captioning. Our model builds on a deep convolutional neural network (CNN) and two separate LSTM networks. It is capable of learning long term visual-language interactions by making use of history and future context information at high level semantic space. Two novel deep bidirectional variant models, in which we increase the depth of nonlinearity transition in different way, are proposed to learn hierarchical visual-language embeddings. Data augmentation techniques such as multi-crop, multi-scale and vertical mirror are proposed to prevent overfitting in training deep models. We visualize the evolution of bidirectional LSTM internal states over time and qualitatively analyze how our models "translate" image to sentence. Our proposed models are evaluated on caption generation and image-sentence retrieval tasks with three benchmark datasets: Flickr8K, Flickr30K and MSCOCO datasets. We demonstrate that bidirectional LSTM models achieve highly competitive performance to the state-of-the-art results on caption generation even without integrating additional mechanism (e.g. object detection, attention model etc.) and significantly outperform recent methods on retrieval task