Stacked Multimodal Attention Network for Context-Aware Video Captioning

Stacked Multimodal Attention Network for Context-Aware Video Captioning
复制标题

DOI:
10.1109/tcsvt.2021.3058626
复制
发表时间:
2021-02
影响因子:
8.4
通讯作者:
Y. Zheng;Yuejie Zhang;Rui Feng;Tao Zhang;Weiguo Fan
Y. Zheng;Yuejie Zhang;Rui Feng;Tao Zhang;Weiguo Fan
中科院分区:
工程技术1区
文献类型:
--
作者:
Y. Zheng;Yuejie Zhang;Rui Feng;Tao Zhang;Weiguo Fan

文献摘要

被引文献

相似文献

最近的视频字幕神经模型通常采用基于注意力的编码器-解码器框架。然而,目前的方法在生成字幕时主要关注视频的运动特征和对象特征,而忽略了潜在但有用的历史信息。此外,现有的字幕生成模型普遍存在曝光偏差和梯度消失问题。在本文中,我们提出了一种新的视频字幕框架,命名为堆叠多模态注意力网络(SMAN)。它采用字幕生成过程中附加的视觉和文本历史信息作为上下文特征,采用堆栈架构逐步处理不同的特征,并利用强化学习方法和由粗到细的训练策略进一步改善生成的结果。在MSVD和MSR-VTT基准数据集上的定量和定性实验均表明了该框架的有效性和可行性。代码可在https://github.com/zhengyi123456/SMAN上获得。
Recent neural models for video captioning usually employ an attention-based encoder-decoder framework. However, current approaches mainly attend to the motion features and object features of the video when generating the caption, but ignore the potential but useful historical information. Besides, exposure bias and vanishing gradients problems always exist in current caption generation models. In this paper, we propose a novel video captioning framework, named Stacked Multimodal Attention Network (SMAN). It adopts additional visual and textual historical information during caption generation as context features, employs a stacked architecture to process different features gradually, and utilizes the Reinforcement Learning method and coarse-to-fine training strategy to further improve the generated results. Both quantitative and qualitative experiments on the benchmark datasets of MSVD and MSR-VTT show the effectiveness and feasibility of our framework. The codes are available on https://github.com/zhengyi123456/SMAN.