QNet: An Adaptive Quantization Table Generator Based on Convolutional Neural Network

QNet: An Adaptive Quantization Table Generator Based on Convolutional Neural Network
复制标题

QNet:基于卷积神经网络的自适应量化表生成器

DOI:
10.1109/tip.2020.3030126
复制
发表时间:
2020-01-01
影响因子:
10.6
通讯作者:
Zeng, Xiaoyang
Zeng, Xiaoyang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yan, Xiao;Fan, Yibo;Zeng, Xiaoyang

文献摘要

被引文献

相似文献

JPEG是使用最广泛的有损图像压缩标准之一,其压缩性能在很大程度上取决于量化表。在这项工作中,我们利用卷积神经网络(CNN)以符合标准的方式生成图像自适应量化表。首先建立一个包含10,000多幅图像的图像集,并通过经典的遗传算法生成其最优量化表,然后提出一种方法,可以有效地提取和融合每幅图像的频域和空域信息,训练回归网络直接生成自适应量化表。此外,我们从数据集中提取了几个有代表性的量化表,并训练了一个分类网络,以指示每个图像的最佳量化表,这进一步提高了压缩性能和计算效率。对不同图像的测试表明,该方法明显优于最先进的方法。在1.0 bpp的压缩率下,与标准表相比,回归和分类网络提供了近1.2和1.4 dB的平均峰值信噪比(PSNR)增益。对于结构相似性指数测量(SSIM)下的实验,改进分别为0.4%和0.54%。所提出的方法也具有竞争力的计算效率,因为回归和分类网络仅需15和6.25毫秒,分别在3.20 GHz的单个CPU核心上处理$768 \times 512$的图像。
The JPEG is one of the most widely used lossy image-compression standards, whose compression performance depends largely on a quantization table. In this work, we utilize a Convolutional Neural Network (CNN) to generate an image-adaptive quantization table in a standard-compliant way. We first build an image set containing more than 10,000 images and generate their optimal quantization tables through a classical genetic algorithm, and then propose a method that can efficiently extract and fuse the frequency and spatial domain information of each image to train a regression network to directly generate adaptive quantization tables. In addition, we extract several representative quantization tables from the dataset and train a classification network to indicate the optimal one for each image, which further improves compression performance and computational efficiency. Tests on diverse images show that the proposed method clearly outperforms the state-of-the-art method. Compared with the standard table at the compression rate of 1.0 bpp, the regression and classification network provide average Peak Signal-to-Noise Ratio (PSNR) gains of nearly 1.2 and 1.4 dB. For the experiment under Structural Similarity Index Measurement (SSIM), the improvements are 0.4% and 0.54%, respectively. The proposed method also has competitive computational efficiency, as the regression and classification network only take 15 and 6.25 milliseconds, respectively, to process a $768 \times 512$ image on a single CPU core at 3.20 GHz.