Action Semantic Alignment for Image Captioning
Action Semantic Alignment for Image Captioning
复制标题
DOI:
10.1109/mipr54900.2022.00041
复制
发表时间:
2022-08
期刊:
影响因子:
--
通讯作者:
Da Huo;Marc A. Kastner;Takahiro Komamizu;I. Ide
中科院分区:
文献类型:
--
作者:
Da Huo;Marc A. Kastner;Takahiro Komamizu;I. Ide
Image captioning is one of the main goals in vision and language processing, which aims to generate proper descriptions of images. Recently, the attention mechanisms became crucial in captioning tasks, as they can capture global dependencies between modalities. Moreover, some works have used objects detected from the input image as anchor points, so called object tags, to ease such alignments resulting in good performance for this task. In this paper, we newly introduce action information as a prior to further improve this, by adding action tags for training. The action tags can learn alignment at action semantic level and catch the previously ignored dimension of action, that could be very important in image captioning. We found that training with action tags can be used to describe images in a dynamic style. Furthermore, we found it can actually lead to a significant improvement compared with other methods in captioning performance measured by common metrics.