Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning

Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning
复制标题

DOI:
10.18653/v1/2020.emnlp-main.61
复制
发表时间:
2020-03
期刊:
--
影响因子:
--
通讯作者:
Zhiyuan Fang;Tejas Gokhale;Pratyay Banerjee;Chitta Baral;Yezhou Yang
Zhiyuan Fang;Tejas Gokhale;Pratyay Banerjee;Chitta Baral;Yezhou Yang
中科院分区:
其他
文献类型:
--
作者:
Zhiyuan Fang;Tejas Gokhale;Pratyay Banerjee;Chitta Baral;Yezhou Yang

文献摘要

被引文献

相似文献

字幕是视频理解中一项重要且具有挑战性的任务。在涉及诸如人类的主动代理的视频中,代理的动作可以在场景中带来无数变化。这些变化是可以观察到的,比如场景中物体的移动、操纵和变换--这些都反映在传统的视频字幕中。然而,与图像不同的是,视频中的动作也与社会和常识方面有内在的联系,例如意图(为什么动作发生),属性(例如谁在做动作,在谁身上,在哪里,使用什么等)。和影响(世界如何因行动而改变,行动对其他代理人的影响)。因此,对于视频理解,例如在为视频添加字幕或回答有关视频的问题时,必须了解这些常识方面。我们提出的第一个工作直接从视频生成\textit{常识}字幕,以描述潜在的方面,如意图,属性和效果。我们提出了一个新的数据集“视频到常识(V2 C)”,其中包含9 k人类代理执行各种动作的视频,注释有3种类型的常识描述。此外,我们探索使用开放式的基于视频的常识问答(V2 C-QA)作为一种方式来丰富我们的字幕。我们在V2 C-QA任务上微调了我们的常识生成模型,我们询问了有关视频中潜在方面的问题。生成任务和QA任务都可以用来丰富视频字幕。
Captioning is a crucial and challenging task for video understanding. In videos that involve active agents such as humans, the agent's actions can bring about myriad changes in the scene. These changes can be observable, such as movements, manipulations, and transformations of the objects in the scene -- these are reflected in conventional video captioning. However, unlike images, actions in videos are also inherently linked to social and commonsense aspects such as intentions (why the action is taking place), attributes (such as who is doing the action, on whom, where, using what etc.) and effects (how the world changes due to the action, the effect of the action on other agents). Thus for video understanding, such as when captioning videos or when answering question about videos, one must have an understanding of these commonsense aspects. We present the first work on generating \textit{commonsense} captions directly from videos, in order to describe latent aspects such as intentions, attributes, and effects. We present a new dataset "Video-to-Commonsense (V2C)" that contains 9k videos of human agents performing various actions, annotated with 3 types of commonsense descriptions. Additionally we explore the use of open-ended video-based commonsense question answering (V2C-QA) as a way to enrich our captions. We finetune our commonsense generation models on the V2C-QA task where we ask questions about the latent aspects in the video. Both the generation task and the QA task can be used to enrich video captions.