TVQA: Localized, Compositional Video Question Answering

TVQA: Localized, Compositional Video Question Answering
复制标题

DOI:
10.18653/v1/d18-1167
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Jie Lei;Licheng Yu;Mohit Bansal;Tamara L. Berg
Jie Lei;Licheng Yu;Mohit Bansal;Tamara L. Berg
中科院分区:
其他
文献类型:
--
作者:
Jie Lei;Licheng Yu;Mohit Bansal;Tamara L. Berg

文献摘要

被引文献

相似文献

近年来,人们对基于图像的问答(QA)任务越来越感兴趣。然而,由于数据的限制,基于视频的QA的工作要少得多。在本文中,我们提出了TVQA,一个大规模的视频问答数据集的基础上,6个流行的电视节目。TVQA由来自21,793个剪辑的152,545个问答对组成,跨越460小时的视频。问题被设计为本质上是合成的,需要系统联合定位剪辑中的相关时刻,理解基于字幕的对话,并识别相关的视觉概念。我们提供了对这个新数据集的分析,以及几个基线和一个用于TVQA任务的多流端到端可训练神经网络框架。该数据集可在http://tvqa.cs.unc.edu上公开获得。
Recent years have witnessed an increasing interest in image-based question-answering (QA) tasks. However, due to data limitations, there has been much less work on video-based QA. In this paper, we present TVQA, a large-scale video QA dataset based on 6 popular TV shows. TVQA consists of 152,545 QA pairs from 21,793 clips, spanning over 460 hours of video. Questions are designed to be compositional in nature, requiring systems to jointly localize relevant moments within a clip, comprehend subtitle-based dialogue, and recognize relevant visual concepts. We provide analyses of this new dataset as well as several baselines and a multi-stream end-to-end trainable neural network framework for the TVQA task. The dataset is publicly available at http://tvqa.cs.unc.edu.