Predicting Human Scanpaths in Visual Question Answering

Predicting Human Scanpaths in Visual Question Answering
复制标题

DOI:
10.1109/cvpr46437.2021.01073
复制
发表时间:
2021-06
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Xianyu Chen;Ming Jiang;Qi Zhao
Xianyu Chen;Ming Jiang;Qi Zhao
中科院分区:
其他
文献类型:
--
作者:
Xianyu Chen;Ming Jiang;Qi Zhao

文献摘要

相似文献

注意力一直是人类和计算机视觉系统的重要机制。虽然预测注意力的最新模型专注于估计具有自由观看行为的静态概率显着性图,但现实生活中的场景充满了不同类型和复杂性的任务,视觉探索是一个有助于任务性能的时间过程。为了弥补这一差距,我们进行了第一项研究,以了解和预测的时间序列的眼睛注视(又名。扫描路径),并研究扫描路径如何影响任务性能。我们提出了一种新的深度强化学习方法来预测在视觉问答中导致不同表现的扫描路径。在任务引导图的条件下,该模型学习特定于问题的注意模式来生成扫描路径。它解决了自我批判序列训练的扫描路径预测中的曝光偏差问题,并设计了一个一致性-发散性损失,以生成正确和不正确答案之间可区分的扫描路径。该模型不仅准确预测了视觉问答中人类行为的时空模式,如固定位置,持续时间和顺序,而且还推广到自由观看和视觉搜索任务,在所有任务中实现了人类水平的性能,并显着优于现有技术。
Attention has been an important mechanism for both humans and computer vision systems. While state-of-the-art models to predict attention focus on estimating a static probabilistic saliency map with free-viewing behavior, real-life scenarios are filled with tasks of varying types and complexities, and visual exploration is a temporal process that contributes to task performance. To bridge the gap, we conduct a first study to understand and predict the temporal sequences of eye fixations (a.k.a. scanpaths) during performing general tasks, and examine how scanpaths affect task performance. We present a new deep reinforcement learning method to predict scanpaths leading to different performances in visual question answering. Conditioned on a task guidance map, the proposed model learns question-specific attention patterns to generate scanpaths. It addresses the exposure bias in scanpath prediction with self-critical sequence training and designs a Consistency-Divergence loss to generate distinguishable scanpaths between correct and incorrect answers. The proposed model not only accurately predicts the spatio-temporal patterns of human behavior in visual question answering, such as fixation position, duration, and order, but also generalizes to free-viewing and visual search tasks, achieving human-level performance in all tasks and significantly outperforming the state of the art.