Temporal Attention and Consistency Measuring for Video Question Answering

Temporal Attention and Consistency Measuring for Video Question Answering
复制标题

DOI:
10.1145/3382507.3418886
复制
发表时间:
2020-10
期刊:
Proceedings of the 2020 International Conference on Multimodal Interaction
影响因子:
--
通讯作者:
Lingyu Zhang;R. Radke
Lingyu Zhang;R. Radke
中科院分区:
其他
文献类型:
--
作者:
Lingyu Zhang;R. Radke

文献摘要

相似文献

社会信号处理算法在解决小组讨论的视听记录中的明确预测和估计问题方面变得越来越好。然而,人类的许多行为和交流是不那么结构化和更微妙的。在本文中,我们解决的问题,从不同的人类互动的视听记录的通用问题回答。我们的目标是为视频中关于人类交互的自由文本问题选择正确的自由文本答案。我们提出了一个基于RNN的模型,其中有两个新颖的想法:一个时间注意力模块,突出问题和候选答案中的关键词和短语,以及一个一致性测量模块,对多模态数据,问题和候选答案之间的相似性进行评分。这一小部分一致性得分构成了最终问答阶段的输入,从而产生了一个轻量级模型。我们证明了我们的模型在包含数百个视频和问题/答案对的Social-IQ数据集上达到了最先进的准确性。
Social signal processing algorithms have become increasingly better at solving well-defined prediction and estimation problems in audiovisual recordings of group discussion. However, much human behavior and communication is less structured and more subtle. In this paper, we address the problem of generic question answering from diverse audiovisual recordings of human interaction. The goal is to select the correct free-text answer to a free-text question about human interaction in a video. We propose an RNN-based model with two novel ideas: a temporal attention module that highlights key words and phrases in the question and candidate answers, and a consistency measurement module that scores the similarity between the multimodal data, the question, and the candidate answers. This small set of consistency scores forms the input to the final question-answering stage, resulting in a lightweight model. We demonstrate that our model achieves state of the art accuracy on the Social-IQ dataset containing hundreds of videos and question/answer pairs.