Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning.

Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning.
复制标题

使用逆增强学习来预测目标指导的人类注意力。

DOI:
10.1109/cvpr42600.2020.00027
复制
发表时间:
2020-06
期刊:
Proceedings. IEEE Computer Society Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Hoai M
Hoai M
中科院分区:
其他
文献类型:
--
作者:
Yang Z;Huang L;Chen Y;Wei Z;Ahn S;Zelinsky G;Samaras D;Hoai M

文献摘要

参考文献

被引文献

相似文献

人类凝视行为预测对于行为视觉和计算机视觉应用具有重要意义。大多数模型主要关注于使用显著图预测自由观看行为,但不能将其推广到目标导向行为,例如当一个人搜索视觉目标对象时。我们提出了第一个反向强化学习(IRL)模型来学习人类在视觉搜索过程中使用的内部奖励函数和策略。我们将观察者的内部信念状态建模为对象位置的动态上下文信念图。这些映射被学习,然后用于预测多个目标类别的行为扫描路径。为了训练和评估我们的IRL模型,我们创建了Coco-Search18,这是目前存在的最大的高质量搜索定位数据集。Coco-Search18有10名参与者在6202张图像中搜索18个目标对象类别中的每一个,进行了大约30万次目标定向凝视。当在Coco-Search18上训练和评估时,IRL模型在预测搜索固定扫描路径方面优于基线模型,无论是在与人类搜索行为的相似性方面还是在搜索效率方面。最后,通过IRL模型恢复的奖励地图揭示了对象优先排序的独特的目标依赖模式,我们将其解释为学习的对象上下文。
Human gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We propose the first inverse reinforcement learning (IRL) model to learn the internal reward function and policy used by humans during visual search. We modeled the viewer’s internal belief states as dynamic contextual belief maps of object locations. These maps were learned and then used to predict behavioral scanpaths for multiple target categories. To train and evaluate our IRL model we created COCO-Search18, which is now the largest dataset of high-quality search fixations in existence. COCO-Search18 has 10 participants searching for each of 18 target-object categories in 6202 images, making about 300,000 goal-directed fixations. When trained and evaluated on COCO-Search18, the IRL model outperformed baseline models in predicting search fixation scanpaths, both in terms of similarity to human search behavior and search efficiency. Finally, reward maps recovered by the IRL model reveal distinctive target-dependent patterns of object prioritization, which we interpret as a learned object context.
DOI: 10.1167/11.5.14
发表时间: 2011-01-01
期刊: JOURNAL OF VISION
影响因子: 1.8
作者:
Eckstein, Miguel P.
通讯作者: Eckstein, Miguel P.
DOI: 10.1109/tpami.2012.89
发表时间: 2013-01-01
影响因子: 23.6
作者:
Borji, Ali;Itti, Laurent
通讯作者: Itti, Laurent
DOI: 10.1167/9.5.19
发表时间: 2009-01-01
期刊: JOURNAL OF VISION
影响因子: 1.8
作者:
Berg, David J.;Boehnke, Susan E.;Itti, Laurent
通讯作者: Itti, Laurent
DOI: 10.1109/34.730558
发表时间: 1998-11-01
影响因子: 23.6
作者:
Itti, L;Koch, C;Niebur, E
通讯作者: Niebur, E
DOI: 10.3758/s13414-014-0764-6
发表时间: 2015-01
期刊: Attention, perception & psychophysics
影响因子: --
作者:
Hout MC;Goldinger SD
通讯作者: Goldinger SD