What/Where to Look Next? Modeling Top-Down Visual Attention in Complex Interactive Environments

What/Where to Look Next? Modeling Top-Down Visual Attention in Complex Interactive Environments
复制标题

DOI:
10.1109/tsmc.2013.2279715
复制
发表时间:
2014-05-01
影响因子:
8.7
通讯作者:
Itti, Laurent
Itti, Laurent
中科院分区:
计算机科学1区
文献类型:
--
作者:
Borji, Ali;Sihite, Dicky N.;Itti, Laurent

文献摘要

被引文献

相似文献

已经提出了几种视觉注意模型,用于描述在简单刺激和任务(如自由观看或视觉搜索)下的眼睛运动。然而,到目前为止,还没有一个计算框架可以可靠地模拟人类在更复杂的环境和任务中的凝视行为,比如城市驾驶。此外,基准数据集、评分技术和自上而下的模型体系结构还没有得到很好的理解。在本文中,我们描述了基于概率推理和推理的图形模型的自上而下显性视觉注意建模的新的任务相关方法。我们描述了一种动态贝叶斯网络,它直接从观测数据推断受监视对象和空间位置的概率分布。在我们的模型中,概率推理是在与对象相关的函数上执行的,这些函数是从视频场景中的对象的手动注释或通过最新的对象检测/识别算法提供的。对玩三种视频游戏(时间安排、驾驶和飞行战斗)的观察者进行了大约3小时(大约315,000次眼睛注视和12 600次扫视)的评估,我们发现我们的方法比以下方法更能预测眼睛注视:1)这里还开发了更简单的基于分类器的模型,将场景的特征(从主旨、自下而上的显著程度、身体动作和事件的多模式信息)映射到眼睛位置;2)14种最先进的自下而上的显著程度模型;以及3)暴力算法,如平均眼睛位置。实验结果表明,与现有模型相比,该模型在时空视觉数据的使用和推理方面更有效。
Several visual attention models have been proposed for describing eye movements over simple stimuli and tasks such as free viewing or visual search. Yet, to date, there exists no computational framework that can reliably mimic human gaze behavior in more complex environments and tasks such as urban driving. In addition, benchmark datasets, scoring techniques, and top-down model architectures are not yet well understood. In this paper, we describe new task-dependent approaches for modeling top-down overt visual attention based on graphical models for probabilistic inference and reasoning. We describe a dynamic Bayesian network that infers probability distributions over attended objects and spatial locations directly from observed data. Probabilistic inference in our model is performed over object-related functions that are fed from manual annotations of objects in video scenes or by state-of-the-art object detection/recognition algorithms. Evaluating over approximately 3 h (approximately 315 000 eye fixations and 12 600 saccades) of observers playing three video games (time-scheduling, driving, and flight combat), we show that our approach is significantly more predictive of eye fixations compared to: 1) simpler classifier-based models also developed here that map a signature of a scene (multimodal information from gist, bottom-up saliency, physical actions, and events) to eye positions; 2) 14 state-of-the-art bottom-up saliency models; and 3) brute-force algorithms such as mean eye position. Our results show that the proposed model is more effective in employing and reasoning over spatio-temporal visual data compared with the state-of-the-art.