Variable Rate Image Compression with Recurrent Neural Networks

Variable Rate Image Compression with Recurrent Neural Networks
复制标题

DOI:
--
复制
发表时间:
2015-11
期刊:
CoRR
影响因子:
--
通讯作者:
G. Toderici;Sean M. O'Malley;S. Hwang;Damien Vincent;David C. Minnen;S. Baluja;Michele Covell;
G. Toderici;Sean M. O'Malley;S. Hwang;Damien Vincent;David C. Minnen;S. Baluja;Michele Covell;
中科院分区:
其他
文献类型:
--
作者:
G. Toderici;Sean M. O'Malley;S. Hwang;Damien Vincent;David C. Minnen;S. Baluja;Michele Covell;

文献摘要

被引文献

相似文献

现在,很大一部分互联网流量是由屏幕相对较小的移动设备的请求驱动的,这些设备通常具有严格的带宽要求。由于这些因素,在初始页面加载过程中传输低分辨率、低字节数的图像预览(缩略图)已经成为现代图形化网站的标准,以提高页面响应性。因此,在现有编解码器的能力之外增加缩略图压缩是当前的研究重点,因为任何字节的节省都将显著增强移动设备用户的体验。为此,我们提出了一个可变速率图像压缩的通用框架和一个基于卷积和反卷积LSTM循环网络的新架构。我们的模型解决了阻碍自编码器神经网络与现有图像压缩算法竞争的主要问题:(1)我们的网络只需要训练一次(不是每张图像),而不考虑输入图像的尺寸和期望的压缩率;(2)我们的网络是渐进的,这意味着发送的比特越多,图像重建就越准确;(3)对于给定的比特数,所提出的架构至少与标准的目的训练的自编码器一样高效。在32美元× 32美元缩略图的大规模基准测试中,我们基于lstm的方法提供了比(无标题)JPEG, JPEG2000和WebP更好的视觉质量,存储大小减少了10%或更多。
A large fraction of Internet traffic is now driven by requests from mobile devices with relatively small screens and often stringent bandwidth requirements. Due to these factors, it has become the norm for modern graphics-heavy websites to transmit low-resolution, low-bytecount image previews (thumbnails) as part of the initial page load process to improve apparent page responsiveness. Increasing thumbnail compression beyond the capabilities of existing codecs is therefore a current research focus, as any byte savings will significantly enhance the experience of mobile device users. Toward this end, we propose a general framework for variable-rate image compression and a novel architecture based on convolutional and deconvolutional LSTM recurrent networks. Our models address the main issues that have prevented autoencoder neural networks from competing with existing image compression algorithms: (1) our networks only need to be trained once (not per-image), regardless of input image dimensions and the desired compression rate; (2) our networks are progressive, meaning that the more bits are sent, the more accurate the image reconstruction; and (3) the proposed architecture is at least as efficient as a standard purpose-trained autoencoder for a given number of bits. On a large-scale benchmark of 32$\times$32 thumbnails, our LSTM-based approaches provide better visual quality than (headerless) JPEG, JPEG2000 and WebP, with a storage size that is reduced by 10% or more.