CLEVR_HYP: A Challenge Dataset and Baselines for Visual Question Answering with Hypothetical Actions over Images

CLEVR_HYP: A Challenge Dataset and Baselines for Visual Question Answering with Hypothetical Actions over Images
复制标题

DOI:
10.18653/v1/2021.naacl-main.289
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Shailaja Keyur Sampat;Akshay Kumar;Yezhou Yang;Chitta Baral
Shailaja Keyur Sampat;Akshay Kumar;Yezhou Yang;Chitta Baral
中科院分区:
其他
文献类型:
--
作者:
Shailaja Keyur Sampat;Akshay Kumar;Yezhou Yang;Chitta Baral

文献摘要

相似文献

现有的视觉问答研究大多局限于图像或视频中显式呈现的信息。在这篇文章中,我们将视觉理解带到一个更高的水平,其中系统被挑战来回答涉及在给定场景中执行特定动作的假设后果的心理模拟的问题。为此,我们基于CLEVR(Johnson et.)制定了一个视觉语言问答任务。等,2017)数据集。然后,我们修改了现有的最好的VQA方法,并为这项任务提出了基线解算器。最后,我们通过提供关于不同架构在图像-文本通道上执行联合推理的能力的见解来推动更好的视觉语言模型的开发。我们的数据集设置脚本和代码将在https://github.com/shailaja183/clevr_hyp.上公开提供
Most existing research on visual question answering (VQA) is limited to information explicitly present in an image or a video. In this paper, we take visual understanding to a higher level where systems are challenged to answer questions that involve mentally simulating the hypothetical consequences of performing specific actions in a given scenario. Towards that end, we formulate a vision-language question answering task based on the CLEVR (Johnson et. al., 2017) dataset. We then modify the best existing VQA methods and propose baseline solvers for this task. Finally, we motivate the development of better vision-language models by providing insights about the capability of diverse architectures to perform joint reasoning over image-text modality. Our dataset setup scripts and codes will be made publicly available at https://github.com/shailaja183/clevr_hyp.