CAVAN: Commonsense Knowledge Anchored Video Captioning

CAVAN: Commonsense Knowledge Anchored Video Captioning
复制标题

DOI:
10.1109/icpr56361.2022.9956241
复制
发表时间:
2022-08
期刊:
2022 26th International Conference on Pattern Recognition (ICPR)
影响因子:
--
通讯作者:
Huiliang Shao;Zhiyuan Fang;Yezhou Yang
Huiliang Shao;Zhiyuan Fang;Yezhou Yang
中科院分区:
其他
文献类型:
--
作者:
Huiliang Shao;Zhiyuan Fang;Yezhou Yang

文献摘要

相似文献

视频剪辑带有的不仅是静态实体的聚集,而且是这些实体之间的各种相互作用和关系。视频字幕系统仍然存在挑战,以生成以突出兴趣并与超出观察结果相符的描述。在这项工作中,我们提出了常识性知识锚定视频字幕(称为Cavan)方法。 Cavan利用推论常识知识,通过新颖的句子级别的语义对齐方式帮助培训视频字幕模型。具体而言,我们通过查询通用知识Atlas(Atomic [1])并形成Commensens optaption optailment Instailment copcus,获得了常识性知识补充每个培训字幕。然后,基于BERT [2]的语言模型从该语料库进行训练,然后用作培训视频字幕模型的常识性歧视器,并惩罚该模型无法生成语义错位的字幕。在MSRVTT [3],V2C [4]和VATEX [5]数据集上进行消融的实验结果验证了Cavan的有效性,并表明使用常识知识有益于视频标题的生成。
It is not merely an aggregation of static entities that a video clip carries, but also a variety of interactions and relations among these entities. Challenges still remain for a video captioning system to generate descriptions focusing on the prominent interest and aligning with the latent aspects beyond observations. In this work, we present a Commonsense knowledge Anchored Video cAptioNing(dubbed as CAVAN) approach. CAVAN exploits inferential commonsense knowledge to assist the training of video captioning model with a novel paradigm for sentence-level semantic alignment. Specifically, we acquire commonsense knowledge complementing per training caption by querying a generic knowledge atlas (ATOMIC [1]), and form the commonsense-caption entailment corpus. A BERT [2] based language entailment model trained from this corpus then serves as a commonsense discriminator for the training of video captioning model, and penalizes the model from generating semantically misaligned captions. Experimental results with ablations on MSRVTT [3], V2C [4] and VATEX [5] datasets validate the effectiveness of CAVAN and reveal that the use of commonsense knowledge benefits video caption generation.