Image Captioning with Bidirectional Semantic Attention-Based Guiding of Long Short-Term Memory.

Image Captioning with Bidirectional Semantic Attention-Based Guiding of Long Short-Term Memory.
复制标题

基于双向语义注意的长短期记忆引导的图像描述

DOI:
10.1007/s11063-018-09973-5
复制
发表时间:
2019-08
影响因子:
3.1
通讯作者:
Guan R
Guan R
中科院分区:
计算机科学4区
文献类型:
--
作者:
Cao P;Yang Z;Sun L;Liang Y;Yang MQ;Guan R

文献摘要

参考文献

被引文献

相似文献

利用自然语言自动描述图像内容,不仅融合了计算机视觉和自然语言处理,而且具有实际应用价值,因此受到了广泛关注。使用端到端的方法,我们提出了一个双向语义注意引导的长短期记忆(Bag-LSTM)模型的图像字幕。该模型有意识地从先前生成的文本中提炼图像特征。通过微调卷积神经网络的参数,Bag-LSTM通过反馈传播获得比其他模型更多的文本相关图像特征。与现有的直接将图像特征添加到LSTM块的每个单元中的指导LSTM方法相反,我们的微调模型动态地利用了更多的文本条件图像特征,这些特征是由语义注意机制获取的,作为指导信息。此外,我们利用双向gLSTM作为字幕生成器,它能够通过利用历史和未来的上下文信息来学习视觉特征和语义信息之间的长期关系。此外,提出了Bag-LSTM模型的变体,以充分描述高级视觉语言交互。在Flickr 8 k和MSCOCO基准数据集上的实验结果表明,该模型在CIDER指标上比BRNN算法提高了51.2%。
Automatically describing contents of an image using natural language has drawn much attention because it not only integrates computer vision and natural language processing but also has practical applications. Using an end-to-end approach, we propose a bidirectional semantic attention-based guiding of long short-term memory (Bag-LSTM) model for image captioning. The proposed model consciously refines image features from previously generated text. By fine-tuning the parameters of convolution neural networks, Bag-LSTM obtains more text-related image features via feedback propagation than other models. As opposed to existing guidance-LSTM methods which directly add image features into each unit of an LSTM block, our fine-tuned model dynamically leverages more text-conditional image features, acquired by the semantic attention mechanism, as guidance information. Moreover, we exploit bidirectional gLSTM as the caption generator, which is capable of learning long term relations between visual features and semantic information by making use of both historical and future contextual information. In addition, variations of the Bag-LSTM model are proposed in an effort to sufficiently describe high-level visual-language interactions. Experiments on the Flickr8k and MSCOCO benchmark datasets demonstrate the effectiveness of the model, as compared with the baseline algorithms, such as it is 51.2% higher than BRNN on CIDEr metric.
DOI: 10.1109/tpami.2012.162
发表时间: 2013-12-01
影响因子: 23.6
作者:
Kulkarni, Girish;Premraj, Visruth;Berg, Tamara L.
通讯作者: Berg, Tamara L.
DOI: 10.1007/978-1-4419-7835-6_10
发表时间: 2011-01-01
期刊: BIOPHYSICAL REGULATION OF VASCULAR DIFFERENTIATION AND ASSEMBLY
影响因子: --
作者:
Diop, Rokhaya;Li, Song
通讯作者: Li, Song