Learned image compression with transformers

Learned image compression with transformers
复制标题

DOI:
10.1117/12.2656516
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Tianma Shen;Y. Liu
Tianma Shen;Y. Liu
中科院分区:
其他
文献类型:
--
作者:
Tianma Shen;Y. Liu

文献摘要

相似文献

近年来,基于深度学习的图像压缩(也称为学习的图像压缩)取得了巨大进步。准确的熵模型对于学习的图像压缩至关重要,因为它可以以较低的比特速率压缩高质量的图像。当前学到的图像压缩方案使用上下文模型和高度培训开发了熵模型。上下文模型利用潜在表示中的局部相关性,以提高概率分布近似值,而Hyperpriors则提供侧面信息以估计分布参数。最近,一些基于变压器的学术图像压缩算法已经出现并实现了最先进的速率失真性能,超过了现有的基于基于的卷积神经网络(CNN)的学习图像压缩和传统的图像压缩。与CNN相比,变形金刚更擅长建模长距离依赖性和提取全局特征。但是,基于变压器的图像压缩的研究仍处于早期阶段。在这项工作中,我们提出了一种新型的基于变压器的学术图像压缩模型。它在主图像编码器和解码器以及上下文模型中采用变压器结构。特别是,我们提出了一个基于变压器的空间通道自动回归上下文模型。编码的潜在空间特征分为空间通道块,熵是在频道的第一顺序中依次编码的,然后是2D Zigzag空间顺序,以先前解码的特征块为条件。为了降低计算复杂性,我们还采用了一个滑动窗口来限制参与熵模型的块数量。关于公共图像压缩数据集的实验研究表明,我们提出的基于变压器的图像编解码器在视觉和定量上优于传统的图像压缩和现有学到的图像压缩模型。
Recent years have witnessed great advances in deep learning-based image compression, also known as learned image compression. An accurate entropy model is essential in learned image compression, since it can compress high-quality images with a lower bit rate. Current learned image compression schemes developed entropy models using context models and hyperpriors. Context models utilize local correlations within latent representations for better probability distribution approximation, while hyperpriors provide side information to estimate distribution parameters. Most recently, several transformer-based learned image compression algorithms have emerged and achieved state-of-the-art rate distortion performances, surpassing existing convolutional neural network (CNN)- based learned image compression and traditional image compression. Transformers are better at modeling long-distance dependencies and extracting global features than CNNs. However, the research of transformer-based image compression is still in its early stage. In this work, we propose a novel transformer-based learned image compression model. It adopts transformer structures in the main image encoder and decoder and in the context model. In particular, we propose a transformer-based spatial-channel auto-regressive context model. Encoded latent-space features are split into spatial-channel chunks, which are entropy encoded sequentially in a channelfirst order, followed by a 2D zigzag spatial order, conditioned on previously decoded feature chunks. To reduce the computational complexity, we also adopt a sliding window to restrict the number of chunks participating in the entropy model. Experimental studies on public image compression datasets demonstrate that our proposed transformer-based learned image codec outperforms traditional image compression and existing learned image compression models visually and quantitatively.