A generative adversarial network for video compression

A generative adversarial network for video compression
复制标题

DOI:
10.1117/12.2618714
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Pengli Du;Ying Liu;Nam Ling;Lingzhi Liu;Yongxiong Ren;M. Hsu
Pengli Du;Ying Liu;Nam Ling;Lingzhi Liu;Yongxiong Ren;M. Hsu
中科院分区:
其他
文献类型:
--
作者:
Pengli Du;Ying Liu;Nam Ling;Lingzhi Liu;Yongxiong Ren;M. Hsu

文献摘要

相似文献

视频数据已经占据了人们日常的专业和娱乐活动。这给互联网带宽带来了很大的压力。因此,重要的是开发有效的视频编码技术以尽可能多地压缩视频数据并节省传输带宽,同时仍然提供视觉上令人愉悦的解码视频。在诸如高效视频编码(HEVC)和通用视频编码(VVC)的传统视频编码中,基于信号处理和信息理论的技术是主流。近年来,由于深度学习的进步,出现了许多基于深度学习的图像和视频压缩方法。特别是,生成对抗网络(GAN)已经显示出上级的图像压缩性能。解码后的图像通常比纯卷积神经网络(CNN)图像压缩更清晰,呈现更多细节,并且与人类视觉系统(HVS)更一致。尽管如此,大多数现有的基于GAN的方法都是用于静态图像压缩的,并且很少有研究调查GAN用于视频压缩的潜力。在这项工作中,我们提出了一种新的帧间视频编码方案,压缩参考帧和目标(残留)帧的GAN。由于剩余信号包含较少的能量,所提出的方法有效地降低了比特率。同时,由于我们采用对抗学习,解码目标帧的感知质量得到了很好的保护。我们提出的算法的有效性证明了常见的测试视频序列的实验研究。
Video data has occupied people’s daily professional and entertainment activities. It imposes a big pressure on the Internet bandwidth. Hence, it is important to develop effective video coding techniques to compress video data as much as possible and save the transmission bandwidth, while still providing visually pleasing decoded videos. In conventional video coding such as the high efficiency video coding (HEVC) and the versatile video coding (VVC), signal processing and information theory-based techniques are mainstream. In recent years, thanks to the advances in deep learning, a lot of deep learning-based approaches have emerged for image and video compression. In particular, the generative adversarial networks (GAN) have shown superior performance for image compression. The decoded images are usually sharper and present more details than pure convolutional neural network (CNN)-based image compression and are more consistent with human visual system (HVS). Nevertheless, most existing GAN-based methods are for still image compression, and truly little research investigates the potential of GAN for video compression. In this work, we propose a novel inter-frame video coding scheme that compresses both reference frames and target (residue) frames by GAN. Since residue signals contain less energy, the proposed method effectively reduces the bit rates. Meanwhile, since we adopt adversarial learning, the perceptual quality of decoded target frames is well-preserved. The effectiveness of our proposed algorithm is demonstrated by experimental studies on common test video sequences.