CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering

CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering
复制标题

DOI:
10.48550/arxiv.2211.03779
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Maitreya Patel;Tejas Gokhale;Chitta Baral;Yezhou Yang
Maitreya Patel;Tejas Gokhale;Chitta Baral;Yezhou Yang
中科院分区:
其他
文献类型:
--
作者:
Maitreya Patel;Tejas Gokhale;Chitta Baral;Yezhou Yang

文献摘要

相似文献

视频通常会捕捉物体、它们的可见属性、它们的运动以及不同物体之间的相互作用。物体也具有诸如质量等物理属性,而成像流程无法直接捕捉这些属性。然而,可以通过利用来自相对物体运动的线索以及碰撞所引入的动力学来估计这些属性。在本文中,我们介绍CRIPP - VQA,这是一个新的视频问答数据集,用于对场景中物体的隐含物理属性进行推理。CRIPP - VQA包含运动物体的视频,并标注了涉及对行为效果进行反事实推理的问题、关于为达到目标而进行规划的问题以及关于物体可见属性的描述性问题。CRIPP - VQA测试集能够在几种分布外设置下进行评估——具有在训练分布中未出现的质量、摩擦系数和初始速度的物体的视频。我们的实验揭示了在回答关于物体隐含属性(本文重点)和显式属性(先前工作重点)的问题方面存在令人惊讶且显著的性能差距。
Videos often capture objects, their visible properties, their motion, and the interactions between different objects. Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture. However, these properties can be estimated by utilizing cues from relative object motion and the dynamics introduced by collisions. In this paper, we introduce CRIPP-VQA, a new video question answering dataset for reasoning about the implicit physical properties of objects in a scene. CRIPP-VQA contains videos of objects in motion, annotated with questions that involve counterfactual reasoning about the effect of actions, questions about planning in order to reach a goal, and descriptive questions about visible properties of objects. The CRIPP-VQA test set enables evaluation under several out-of-distribution settings – videos with objects with masses, coefficients of friction, and initial velocities that are not observed in the training distribution. Our experiments reveal a surprising and significant performance gap in terms of answering questions about implicit properties (the focus of this paper) and explicit properties of objects (the focus of prior work).