Benchmarking Gaze Prediction for Categorical Visual Search

Benchmarking Gaze Prediction for Categorical Visual Search
复制标题

DOI:
10.1109/cvprw.2019.00111
复制
发表时间:
2019-06
期刊:
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
G. Zelinsky;Zhibo Yang;Lihan Huang;Yupei Chen;Seoyoung Ahn;Zijun Wei;Hossein Adeli;D. Samaras;Minh Hoai
G. Zelinsky;Zhibo Yang;Lihan Huang;Yupei Chen;Seoyoung Ahn;Zijun Wei;Hossein Adeli;D. Samaras;Minh Hoai
中科院分区:
其他
文献类型:
--
作者:
G. Zelinsky;Zhibo Yang;Lihan Huang;Yupei Chen;Seoyoung Ahn;Zijun Wei;Hossein Adeli;D. Samaras;Minh Hoai

文献摘要

被引文献

相似文献

人类注意力转移的预测是行为和计算机视觉中广泛研究的问题,特别是在自由观看任务的背景下。然而,搜索行为,其中固定扫描路径高度依赖于观众的目标,已经收到了少得多的关注,即使视觉搜索构成了一个人的日常行为。其中一个原因是缺乏可以训练搜索模型的真实图像数据集。在本文中,我们提出了一个精心创建的数据集,用于两个目标类别,微波和时钟,从COCO 2014数据集策划。总共有2183张图像被呈现给多个参与者,他们的任务是搜索两个类别中的一个。这总共产生了16184个用于训练的有效注视点,使我们的微波时钟数据集成为目前分类搜索中最大的眼睛注视点数据集之一。我们还提出了一个40图像测试数据集,其中图像描绘了微波和时钟目标。不同的固定模式取决于参与者是否在相同的图像中搜索微波(n=30)或时钟(n=30),这意味着模型需要从相同的像素输入预测不同的搜索扫描路径。我们报告了在这些数据集上训练和评估的几个最先进的深度网络模型的结果。总的来说,这些数据集和我们的评估协议提供了一个有用的测试平台,用于开发预测特定类别视觉搜索行为的新方法。
The prediction of human shifts of attention is a widelystudied question in both behavioral and computer vision, especially in the context of a free viewing task. However, search behavior, where the fixation scanpaths are highly dependent on the viewer’s goals, has received far less attention, even though visual search constitutes much of a person’s everyday behavior. One reason for this is the absence of real-world image datasets on which search models can be trained. In this paper we present a carefully created dataset for two target categories, microwaves and clocks, curated from the COCO2014 dataset. A total of 2183 images were presented to multiple participants, who were tasked to search for one of the two categories. This yields a total of 16184 validated fixations used for training, making our microwave-clock dataset currently one of the largest datasets of eye fixations in categorical search. We also present a 40-image testing dataset, where images depict both a microwave and a clock target. Distinct fixation patterns emerged depending on whether participants searched for a microwave (n=30) or a clock (n=30) in the same images, meaning that models need to predict different search scanpaths from the same pixel inputs. We report the results of several state-of-the-art deep network models that were trained and evaluated on these datasets. Collectively, these datasets and our protocol for evaluation provide what we hope will be a useful test-bed for the development of new methods for predicting category-specific visual search behavior.