A coherent computational approach to model bottom-up visual attention

A coherent computational approach to model bottom-up visual attention
复制标题

DOI:
10.1109/tpami.2006.86
复制
发表时间:
2006-05-01
影响因子:
23.6
通讯作者:
Thoreau, D
Thoreau, D
中科院分区:
计算机科学1区
文献类型:
--
作者:
Le Meur, O;Le Callet, P;Thoreau, D

文献摘要

被引文献

相似文献

视觉注意力是一种过滤掉冗余视觉信息并检测我们视野中最相关部分的机制。自动确定最视觉相关的区域将在许多应用中是有用的,例如图像和视频编码、水印、视频浏览和质量评估。许多研究小组目前正在研究视觉注意系统的计算建模。第一批发布的计算模型基于一些基本且已广为人知的人类视觉系统(HVS)属性。这些模型具有单一的感知层,仅模拟视觉系统的一个方面。最近的模型集成了HVS的复杂功能,并模拟了视觉输入的分层感知表示。自下而上的机制是现代模型中最常见的特征。这种机制是指非自愿注意(即,不费力或不自觉地吸引我们注意力的显著空间视觉特征)。本文提出了一种自下而上的视觉注意建模的相干计算方法。该模型主要基于目前对HVS行为的理解。对比敏感度函数、感知分解、视觉掩蔽和中心-环绕交互是该模型中实现的一些功能。该算法的性能进行了评估,使用自然图像和实验测量的眼睛跟踪系统。两个适当的众所周知的度量(相关系数和Kullbacl-Leibler分歧)被用来验证这个模型。还定义了另一个度量。最后,将该模型的结果与参考自下而上模型的结果进行比较。
Visual attention is a mechanism which filters out redundant visual information and detects the most relevant parts of our visual field. Automatic determination of the most visually relevant areas would be useful in many applications such as image and video coding, watermarking, video browsing, and quality assessment. Many research groups are currently investigating computational modeling of the visual attention system. The first published computational models have been based on some basic and well-understood Human Visual System (HVS) properties. These models feature a single perceptual layer that simulates only one aspect of the visual system. More recent models integrate complex features of the HVS and simulate hierarchical perceptual representation of the visual input. The bottom-up mechanism is the most occurring feature found in modern models. This mechanism refers to involuntary attention (i.e., salient spatial visual features that effortlessly or involuntary attract our attention). This paper presents a coherent computational approach to the modeling of the bottom-up visual attention. This model is mainly based on the current understanding of the HVS behavior. Contrast sensitivity functions, perceptual decomposition, visual masking, and center-surround interactions are some of the features implemented in this model. The performances of this algorithm are assessed by using natural images and experimental measurements from an eye-tracking system. Two adequate well-known metrics (correlation coefficient and Kullbacl-Leibler divergence) are used to validate this model. A further metric is also defined. The results from this model are finally compared to those from a reference bottom-up model.