How to Make a BLT Sandwich? Learning VQA towards Understanding Web Instructional Videos

How to Make a BLT Sandwich? Learning VQA towards Understanding Web Instructional Videos
复制标题

DOI:
10.1109/wacv48630.2021.00117
复制
发表时间:
2021-01
期刊:
2021 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Shaojie Wang;Wentian Zhao;Ziyi Kou;Jing Shi;Chenliang Xu
Shaojie Wang;Wentian Zhao;Ziyi Kou;Jing Shi;Chenliang Xu
中科院分区:
其他
文献类型:
--
作者:
Shaojie Wang;Wentian Zhao;Ziyi Kou;Jing Shi;Chenliang Xu

文献摘要

相似文献

网络教学视频的理解是视频理解的一个重要分支。首先,大多数现有的视频方法集中在几秒钟长的视频剪辑的短期动作;这些方法不直接适用于长视频。其次,与不受约束的长视频不同,例如,电影、教学视频更有结构性,因为它们具有限制理解任务的逐步过程。在这项工作中,我们研究了问题解决的教学视频通过视觉提问(VQA)。令人惊讶的是,尽管它有丰富的应用程序,但它并没有成为视频社区的重点。因此,我们引入了YouCookQA,这是一个基于YouCook 2的教学视频注释QA数据集[27]。YouCookQA中的问题不限于单个帧上的线索,而是时间维度上多个帧之间的关系。观察缺乏有效的表示建模长视频,我们提出了一套精心设计的模型,包括一个递归图卷积网络(RGCN),捕捉时间顺序和关系信息。此外,我们研究了多种形式,包括描述和成绩单,以提高视频理解的目的。YouCookQA上的大量实验表明,RGCN在QA准确性方面表现最好,并且通过引入人工注释的描述获得了更好的性能。YouCookQA数据集可在https://github.com/Jossome/YoucookQA上获得。
Understanding web instructional videos is an essential branch of video understanding in two aspects. First, most existing video methods focus on short-term actions for a-few-second-long video clips; these methods are not directly applicable to long videos. Second, unlike unconstrained long videos, e.g., movies, instructional videos are more structured in that they have step-by-step procedures constraining the understanding task. In this work, we study problem-solving on instructional videos via Visual Question Answering (VQA). Surprisingly, it has not been an emphasis for the video community despite its rich applications. We thereby introduce YouCookQA, an annotated QA dataset for instructional videos based on YouCook2 [27]. The questions in YouCookQA are not limited to cues on a single frame but relations among multiple frames in the temporal dimension. Observing the lack of effective representations for modeling long videos, we propose a set of carefully designed models including a Recurrent Graph Convolutional Network (RGCN) that captures both temporal order and relational information. Furthermore, we study multiple modalities including descriptions and transcripts for the purpose of boosting video understanding. Extensive experiments on YouCookQA suggest that RGCN performs the best in terms of QA accuracy and better performance is gained by introducing human-annotated descriptions. YouCookQA dataset is available at https://github.com/Jossome/YoucookQA.