Ieee Transactions on Systems Man and Cybernetics Part A-systems and Humans 1 What/where to Look Next? Modeling Top-down Visual Attention in Complex Interactive Environments

Ieee Transactions on Systems Man and Cybernetics Part A-systems and Humans 1 What/where to Look Next? Modeling Top-down Visual Attention in Complex Interactive Environments
复制标题

DOI:
--
复制
发表时间:
--
期刊:
--
影响因子:
--
通讯作者:
A. Borji;Dicky N. Sihite;L. Itti
A. Borji;Dicky N. Sihite;L. Itti
中科院分区:
其他
文献类型:
--
作者:
A. Borji;Dicky N. Sihite;L. Itti

文献摘要

被引文献

相似文献

-已经提出了几个视觉注意模型来描述简单刺激和任务(如自由观看或视觉搜索)的眼球运动。然而到目前为止,还没有一个计算框架可以可靠地模拟人类在更复杂的环境和任务(如城市驾驶)中的凝视行为。此外,基准数据集、评分技术和自顶向下的模型体系结构还没有得到很好的理解。在这项研究中,我们描述了基于概率推理和推理的图形模型的自上而下的显性视觉注意建模的新任务依赖方法。我们描述了一个动态贝叶斯网络(DBN),它可以直接从观测数据推断出被关注对象和空间位置的概率分布。我们模型中的概率推理是在对象相关函数上执行的,这些函数是由视频场景中对象的手动注释或最先进的对象检测/识别算法提供的。在~ 3小时内评估(appx)。315,000次眼睛注视和12,600次扫视)的观察者玩3种电子游戏(时间安排,驾驶和飞行战斗),我们表明,与以下方法相比,我们的方法明显更能预测眼睛注视:(1)这里还开发了更简单的基于分类器的模型,将场景的特征(来自gist、自下而上显著性、物理动作和事件的多模态信息)映射到眼睛位置;(2)14个最先进的自下而上显著性模型;(3)暴力算法,如平均眼睛位置。研究结果表明,与现有模型相比,该模型在时空视觉数据的应用和推理方面更有效。
—Several visual attention models have been proposed for describing eye movements over simple stimuli and tasks such as free viewing or visual search. Yet to date, there exists no computational framework that can reliably mimic human gaze behavior in more complex environments and tasks such as urban driving. Additionally, benchmark datasets, scoring techniques, and top-down model architectures are not yet well understood. In this study, we describe new task-dependent approaches for modeling top-down overt visual attention based on graphical models for probabilistic inference and reasoning. We describe a Dynamic Bayesian Network (DBN) that infers probability distributions over attended objects and spatial locations directly from observed data. Probabilistic inference in our model is performed over object-related functions which are fed from manual annotations of objects in video scenes or by state-of-the-art object detection/recognition algorithms. Evaluating over ∼3 hours (appx. 315, 000 eye fixations and 12, 600 saccades) of observers playing 3 video games (time-scheduling, driving, and flight combat), we show that our approach is significantly more predictive of eye fixations compared to: (1) simpler classifier-based models also developed here that map a signature of a scene (multi-modal information from gist, bottom-up saliency, physical actions, and events) to eye positions, (2) 14 state-of-the-art bottom-up saliency models, and (3) brute-force algorithms such as mean eye position. Our results show that the proposed model is more effective in employing and reasoning over spatio-temporal visual data compared with the state-of-the-art.