A-CAP: Anticipation Captioning with Commonsense Knowledge

A-CAP: Anticipation Captioning with Commonsense Knowledge
复制标题

DOI:
10.1109/cvpr52729.2023.01042
复制
发表时间:
2023-04
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
D. Vo;Quoc-An Luong;Akihiro Sugimoto;Hideki Nakayama
D. Vo;Quoc-An Luong;Akihiro Sugimoto;Hideki Nakayama
中科院分区:
其他
文献类型:
--
作者:
D. Vo;Quoc-An Luong;Akihiro Sugimoto;Hideki Nakayama

文献摘要

相似文献

根据随着时间的推移获得的视觉提示集合的稀疏集合,人类具有推理未来的能力。为了模仿这种能力,我们介绍了一个名为“预期”字幕的新任务,该任务使用漫长的时间订购的图像集为看不见的甲骨文图像生成标题。为了解决这项新任务,我们提出了一个称为A-CAP的模型,该模型将常识性知识纳入了预训练的视觉语言模型中,从而可以预测标题。通过定性和定量评估在定制的视觉讲故事数据集上,A-CAP效果超过其他图像字幕方法,并为预期字幕建立了强大的基线。我们还解决了此任务固有的挑战。
Humans possess the capacity to reason about the future based on a sparse collection of visual cues acquired over time. In order to emulate this ability, we introduce a novel task called Anticipation Captioning, which generates a caption for an unseen oracle image using a sparsely temporally-ordered set of images. To tackle this new task, we propose a model called A-CAP, which incorporates commonsense knowledge into a pre-trained vision-language model, allowing it to anticipate the caption. Through both qualitative and quantitative evaluations on a customized visual storytelling dataset, A-CAP out-performs other image captioning methods and establishes a strong baseline for anticipation captioning. We also address the challenges inherent in this task.