Neural Inter-Frame Compression for Video Coding

Neural Inter-Frame Compression for Video Coding
复制标题

DOI:
10.1109/iccv.2019.00652
复制
发表时间:
2019-10
期刊:
2019 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Abdelaziz Djelouah;Joaquim Campos;Simone Schaub-Meyer;Christopher Schroers
Abdelaziz Djelouah;Joaquim Campos;Simone Schaub-Meyer;Christopher Schroers
中科院分区:
其他
文献类型:
--
作者:
Abdelaziz Djelouah;Joaquim Campos;Simone Schaub-Meyer;Christopher Schroers

文献摘要

被引文献

相似文献

虽然有许多基于深度学习的单图像压缩方法,但端到端学习视频编码领域的探索仍然少得多。因此,在这项工作中,我们提出了一种用于神经视频编码的帧间压缩方法,可以无缝地建立在不同的现有神经图像编解码器上。我们的端到端解决方案通过基于光流的运动补偿在像素空间中执行时间预测。关键的见解是,我们可以通过将所需信息编码到直接解码为运动和混合系数的潜在表示中来提高解码效率和重建质量。为了考虑剩余的预测误差,需要原始图像和内插帧之间的残差信息。我们建议直接在潜在空间而不是像素空间中计算残差,因为这允许对关键帧和中间帧重用相同的图像压缩网络。我们对不同数据集和分辨率的扩展评估表明,我们的方法的率失真性能与现有的最先进的编解码器具有竞争力。
While there are many deep learning based approaches for single image compression, the field of end-to-end learned video coding has remained much less explored. Therefore, in this work we present an inter-frame compression approach for neural video coding that can seamlessly build up on different existing neural image codecs. Our end-to-end solution performs temporal prediction by optical flow based motion compensation in pixel space. The key insight is that we can increase both decoding efficiency and reconstruction quality by encoding the required information into a latent representation that directly decodes into motion and blending coefficients. In order to account for remaining prediction errors, residual information between the original image and the interpolated frame is needed. We propose to compute residuals directly in latent space instead of in pixel space as this allows to reuse the same image compression network for both key frames and intermediate frames. Our extended evaluation on different datasets and resolutions shows that the rate-distortion performance of our approach is competitive with existing state-of-the-art codecs.