Action Semantic Alignment for Image Captioning

Action Semantic Alignment for Image Captioning
复制标题

DOI:
10.1109/mipr54900.2022.00041
复制
发表时间:
2022-08
期刊:
2022 IEEE 5th International Conference on Multimedia Information Processing and Retrieval (MIPR)
影响因子:
--
通讯作者:
Da Huo;Marc A. Kastner;Takahiro Komamizu;I. Ide
Da Huo;Marc A. Kastner;Takahiro Komamizu;I. Ide
中科院分区:
其他
文献类型:
--
作者:
Da Huo;Marc A. Kastner;Takahiro Komamizu;I. Ide

文献摘要

相似文献

图像字幕是视觉和语言处理的主要目标之一,其目的是生成正确的图像描述。最近,注意机制在字幕任务中变得至关重要,因为它们可以捕获模式之间的全局依赖关系。此外,一些作品使用从输入图像中检测到的对象作为锚点,即所谓的对象标签,以简化这种对齐,从而使该任务具有良好的性能。在本文中,我们引入动作信息作为先验,通过添加动作标签进行训练来进一步改进这一点。动作标签可以在动作语义层面学习对齐,并捕获先前被忽略的动作维度,这在图像字幕中可能非常重要。我们发现使用动作标签的训练可以用来描述动态风格的图像。此外,我们发现,与其他方法相比,它实际上可以显著改善用普通指标衡量的字幕性能。
Image captioning is one of the main goals in vision and language processing, which aims to generate proper descriptions of images. Recently, the attention mechanisms became crucial in captioning tasks, as they can capture global dependencies between modalities. Moreover, some works have used objects detected from the input image as anchor points, so called object tags, to ease such alignments resulting in good performance for this task. In this paper, we newly introduce action information as a prior to further improve this, by adding action tags for training. The action tags can learn alignment at action semantic level and catch the previously ignored dimension of action, that could be very important in image captioning. We found that training with action tags can be used to describe images in a dynamic style. Furthermore, we found it can actually lead to a significant improvement compared with other methods in captioning performance measured by common metrics.