An $\ell_{1/2}$ -Norm Regularizer-Based Sparse Coding Framework for Gaze Prediction in First-Person Videos

An $\ell_{1/2}$ -Norm Regularizer-Based Sparse Coding Framework for Gaze Prediction in First-Person Videos
复制标题

DOI:
10.1109/access.2019.2908010
复制
发表时间:
2019-03
期刊:
影响因子:
3.9
通讯作者:
Yujie Li;Zhenni Li;Atsunori Kanemura
Yujie Li;Zhenni Li;Atsunori Kanemura
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yujie Li;Zhenni Li;Atsunori Kanemura

文献摘要

被引文献

相似文献

预测人类注视对于有效地处理和理解来自第一人称视频(FPV)的大量传入视觉信息非常重要。即使人们在噪声环境中持续注视,大多数现有的注视预测算法都是基于显著性映射,这对真实的世界中的噪声环境是敏感的。基于稀疏性的显著性检测算法与最先进的方法相比表现良好。在本文中,我们应用一种新的显着性检测方法的基础上稀疏编码的$\ell _{1/2}$ -范数预测人类的视线FPV。首先通过超像素提取图像边界作为字典的基础,从中构建稀疏表示模型。对于每个超像素,我们首先计算稀疏重建误差。然后,基于重构误差来更新显著图。为了得到稀疏重构误差,最广泛使用的稀疏约束是$\ell _{1}$ -范数。然而,$\ell _{1}$ -范数会导致稀疏向量中大分量的过度惩罚。我们采用$\ell _{1/2}$ -范数进行稀疏编码,这可以导致比$\ell _{1}$ -范数更精确的凝视预测的更稀疏的解决方案。我们将具有$\ell _{1/2}$ -范数的稀疏编码的复杂非凸优化问题转化为若干个一维极小化问题。通过这种方法,我们可以有效地获得封闭形式的解决方案。使用真实世界的凝视数据集的实验结果表明,该算法的性能优于国家的最先进的FPV的凝视预测方法。
Predicting human gaze is important for efficiently processing and understanding numerous incoming visual information from first-person videos (FPVs). Even though people continuously gaze in noisy environments, most existing gaze prediction algorithms are based on saliency mapping, which is sensitive to noisy surroundings in the real world. Sparsity-based saliency detection algorithms perform favorably against state-of-the-art methods. In this paper, we apply a novel saliency detection method based on sparse coding with the $\ell _{1/2}$ -norm for predicting human gaze in FPVs. Image boundaries are first extracted via superpixels as bases for a dictionary, from which a sparse representation model is constructed. For each superpixel, we first compute sparse reconstruction errors. Then, a saliency map is updated based on the reconstruction errors. To receive the sparse reconstruction errors, the most widely utilized sparse constraint is the $\ell _{1}$ -norm. However, the $\ell _{1}$ -norm leads to over-penalization of large components in a sparse vector. We employ the $\ell _{1/2}$ -norm for sparse coding, which can lead to a sparser solution for a more accurate gaze prediction than the $\ell _{1}$ -norm. We transform the complex nonconvex optimization of sparse coding with the $\ell _{1/2}$ -norm to a number of one-dimensional minimization problems. In this way, we obtain the closed-form solutions efficiently. The experimental results using a real-world gaze dataset demonstrate that the proposed algorithm performs better than the state-of-the-art methods of gaze prediction for FPVs.