A Novel Multiresolution Spatiotemporal Saliency Detection Model and Its Applications in Image and Video Compression

A Novel Multiresolution Spatiotemporal Saliency Detection Model and Its Applications in Image and Video Compression
复制标题

DOI:
10.1109/tip.2009.2030969
复制
发表时间:
2010-01-01
影响因子:
10.6
通讯作者:
Zhang, Liming
Zhang, Liming
中科院分区:
计算机科学1区
文献类型:
--
作者:
Guo, Chenlei;Zhang, Liming

文献摘要

被引文献

相似文献

自然场景中的显著区域通常被认为是人眼通常会聚焦的区域,找到这些区域是对象检测的关键步骤。在计算机视觉中,已经提出了许多模型来模拟眼睛的行为,例如显着性工具箱(STB),神经形态视觉工具箱(NVT)等,但它们需要高计算成本,并且计算有用的结果主要依赖于它们的参数选择。虽然提出了一些基于区域的方法来降低特征图的计算复杂度,但这些方法仍然不能在真实的时间内工作。最近,一种简单快速的方法被称为频谱残差(SR),它使用幅度谱的SR来计算图像的显著图。然而,在我们以前的工作中,我们指出,这是相位谱,而不是幅度谱,图像的傅里叶变换是关键的计算的显著区域的位置,并提出了相位谱傅里叶变换(PFT)模型。在本文中,我们提出了一个四元数表示的图像,这是由强度,颜色和运动特征。基于四元数傅里叶变换的基本原理,提出了一种新的多分辨率时空显著性检测模型--四元数傅里叶变换相位谱(PQFT),通过四元数表示计算图像的时空显著性图。与其他模型不同的是,增加的运动维度允许相位谱表示时空显著性,以便不仅对图像而且对视频执行注意力选择。此外,PQFT模型可以计算从粗到细的各种分辨率下的图像的显著图。因此,层次选择性(HS)框架的PQFT模型的基础上,在这里被引入到构建一个图像的树结构表示。为了提高图像和视频压缩的编码效率,本文提出了一种基于HS的多分辨率小波域视觉聚焦模型(MWDF)。视频,自然图像和心理模式的广泛测试表明,所提出的PQFT模型是更有效的显着性检测,可以预测眼睛注视比其他国家的最先进的模型在以前的文献中。此外,我们的模型需要较低的计算成本,因此,可以在真实的时间。对图像和视频的压缩实验表明,HS-MWDF模型比传统模型具有更高的压缩比。
Salient areas in natural scenes are generally regarded as areas which the human eye will typically focus on, and finding these areas is the key step in object detection. In computer vision, many models have been proposed to simulate the behavior of eyes such as SaliencyToolBox (STB), Neuromorphic Vision Toolkit (NVT), and others, but they demand high computational cost and computing useful results mostly relies on their choice of parameters. Although some region-based approaches were proposed to reduce the computational complexity of feature maps, these approaches still were not able to work in real time. Recently, a simple and fast approach called spectral residual (SR) was proposed, which uses the SR of the amplitude spectrum to calculate the image's saliency map. However, in our previous work, we pointed out that it is the phase spectrum, not the amplitude spectrum, of an image's Fourier transform that is key to calculating the location of salient areas, and proposed the phase spectrum of Fourier transform (PFT) model. In this paper, we present a quaternion representation of an image which is composed of intensity, color, and motion features. Based on the principle of PFT, a novel multiresolution spatiotemporal saliency detection model called phase spectrum of quaternion Fourier transform (PQFT) is proposed in this paper to calculate the spatiotemporal saliency map of an image by its quaternion representation. Distinct from other models, the added motion dimension allows the phase spectrum to represent spatiotemporal saliency in order to perform attention selection not only for images but also for videos. In addition, the PQFT model can compute the saliency map of an image under various resolutions from coarse to fine. Therefore, the hierarchical selectivity (HS) framework based on the PQFT model is introduced here to construct the tree structure representation of an image. With the help of HS, a model called multiresolution wavelet domain foveation (MWDF) is proposed in this paper to improve coding efficiency in image and video compression. Extensive tests of videos, natural images, and psychological patterns show that the proposed PQFT model is more effective in saliency detection and can predict eye fixations better than other state-of-the-art models in previous literature. Moreover, our model requires low computational cost and, therefore, can work in real time. Additional experiments on image and video compression show that the HS-MWDF model can achieve higher compression rate than the traditional model.