SG-FCN: A Motion and Memory-Based Deep Learning Model for Video Saliency Detection

SG-FCN: A Motion and Memory-Based Deep Learning Model for Video Saliency Detection
复制标题

DOI:
10.1109/tcyb.2018.2832053
复制
发表时间:
2018-09
影响因子:
11.8
通讯作者:
Meijun Sun;Ziqi Zhou;Q. Hu;Zheng Wang;Jianmin Jiang
Meijun Sun;Ziqi Zhou;Q. Hu;Zheng Wang;Jianmin Jiang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Meijun Sun;Ziqi Zhou;Q. Hu;Zheng Wang;Jianmin Jiang

文献摘要

被引文献

相似文献

由于将卷积神经网络应用于眼睛注视检测,数据驱动的显著性检测引起了人们的强烈兴趣。虽然已经提出了许多基于图像的显著对象和注视点检测模型,但视频注视点检测仍然需要更多的探索。与图像分析不同,运动和时间信息是影响人们观看视频序列时注意力的重要因素。虽然现有的模型的基础上的局部对比度和低层次的功能已经得到了广泛的研究,他们未能同时考虑帧间运动和相邻视频帧的时间信息,导致处理复杂的场景时,性能不尽如人意。为此,我们提出了一种新的和有效的视频眼睛注视检测模型,以提高显着性检测性能。通过模拟人类在观看视频时的记忆机制和视觉注意机制,提出了一种步长增益的全卷积网络,将时间轴上的记忆信息与空间轴上的运动信息相结合,同时存储当前帧的显著性信息。该模型通过分层训练得到,保证了检测的准确性。与11种最先进的方法进行了广泛的实验比较,结果表明,我们提出的模型在许多公开可用的数据集上优于所有11种方法。
Data-driven saliency detection has attracted strong interest as a result of applying convolutional neural networks to the detection of eye fixations. Although a number of image-based salient object and fixation detection models have been proposed, video fixation detection still requires more exploration. Different from image analysis, motion and temporal information is a crucial factor affecting human attention when viewing video sequences. Although existing models based on local contrast and low-level features have been extensively researched, they failed to simultaneously consider interframe motion and temporal information across neighboring video frames, leading to unsatisfactory performance when handling complex scenes. To this end, we propose a novel and efficient video eye fixation detection model to improve the saliency detection performance. By simulating the memory mechanism and visual attention mechanism of human beings when watching a video, we propose a step-gained fully convolutional network by combining the memory information on the time axis with the motion information on the space axis while storing the saliency information of the current frame. The model is obtained through hierarchical training, which ensures the accuracy of the detection. Extensive experiments in comparison with 11 state-of-the-art methods are carried out, and the results show that our proposed model outperforms all 11 methods across a number of publicly available datasets.