Fixation bank: Learning to reweight fixation candidates

Fixation bank: Learning to reweight fixation candidates
复制标题

DOI:
10.1109/cvpr.2015.7298937
复制
发表时间:
2015-06
期刊:
2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Jiaping Zhao;Christian Siagian;L. Itti
Jiaping Zhao;Christian Siagian;L. Itti
中科院分区:
其他
文献类型:
--
作者:
Jiaping Zhao;Christian Siagian;L. Itti

文献摘要

被引文献

相似文献

预测人类将在场景中凝视的位置有很多实际应用。生物启发的显著模型将视觉刺激分解成多个尺度上的特征地图,然后将不同的特征通道集成在例如线性、最大或地图中。然而,到目前为止,还没有一个被普遍接受的功能集成机制。在这里,我们提出了一种新的数据驱动的解决方案:我们首先通过挖掘训练样本来建立“注视库”,该训练样本保持了给定位置周围4个特征通道(颜色、强度、方向、运动)中的局部激活模式与该位置对应的人类注视密度之间的关联。在测试过程中,我们将特征映射分解成斑点,提取每个斑点周围的局部激活模式,将这些模式与固定组逐个套索进行匹配,并根据重建误差确定斑点的权重。我们最终的显著图是所有斑点的加权和。因此,我们的系统将一些空间和特征上下文信息合并到位置相关的加权机制中。在两个标准数据集(DIEM用于训练和测试,CRCNS仅用于测试;总共23,670个训练和15,793+4,505个测试帧)上进行测试,我们的模型略高于但显著优于7个最新的显著模型。
Predicting where humans will fixate in a scene has many practical applications. Biologically-inspired saliency models decompose visual stimuli into feature maps across multiple scales, and then integrate different feature channels, e.g., in a linear, MAX, or MAP. However, to date there is no universally accepted feature integration mechanism. Here, we propose a new a data-driven solution: We first build a “fixation bank” by mining training samples, which maintains the association between local patterns of activation, in 4 feature channels (color, intensity, orientation, motion) around a given location, and corresponding human fixation density at that location. During testing, we decompose feature maps into blobs, extract local activation patterns around each blob, match those patterns against the fixation bank by group lasso, and determine weights of blobs based on reconstruction errors. Our final saliency map is the weighted sum of all blobs. Our system thus incorporates some amount of spatial and featural context information into the location-dependent weighting mechanism. Tested on two standard data sets (DIEM for training and test, and CRCNS for test only; total of 23,670 training and 15,793 + 4,505 test frames), our model slightly but significantly outperforms 7 state-of-the-art saliency models.