Image Coding For Machines: an End-To-End Learned Approach

Image Coding For Machines: an End-To-End Learned Approach
复制标题

机器图像编码:一种端到端的学习方法

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Esa Rahtu
Esa Rahtu
中科院分区:
--
文献类型:
--
作者:
N. Le;Honglei Zhang;Francesco Cricri;R. G. Youvalari;Esa Rahtu

文献摘要

被引文献

相似文献

近年来,基于深度学习的计算机视觉系统以越来越快的速度应用于图像,通常代表了这些图像的唯一消费类型。考虑到每天生成的图像数量急剧增加,一个问题出现了:针对机器消费的图像编解码器与针对人类消费的最先进编解码器相比,性能会好多少?本文提出了一种基于神经网络的端到端学习机器图像编解码器。特别是,我们提出了一套训练策略,以解决平衡竞争损失函数的微妙问题,如计算机视觉任务损失、图像失真损失和速率损失。我们的实验结果表明,我们基于nn的编解码器在目标检测和实例分割任务上优于最先进的Versa-tile视频编码(VVC)标准,分别实现了-37.87%和-32.90%的BD-rate增益,同时由于其紧凑的尺寸而速度很快。据我们所知,这是第一个端到端学习机器目标图像编解码器。
Over recent years, deep learning-based computer vision systems have been applied to images at an ever-increasing pace, oftentimes representing the only type of consumption for those images. Given the dramatic explosion in the number of images generated per day, a question arises: how much better would an image codec targeting machine-consumption perform against state-of-the-art codecs targeting human-consumption? In this paper, we propose an image codec for machines which is neural network (NN) based and end-to-end learned. In particular, we propose a set of training strategies that address the delicate problem of balancing competing loss functions, such as computer vision task losses, image distortion losses, and rate loss. Our experimental results show that our NN-based codec outperforms the state-of-the-art Versa-tile Video Coding (VVC) standard on the object detection and instance segmentation tasks, achieving -37.87% and -32.90% of BD-rate gain, respectively, while being fast thanks to its compact size. To the best of our knowledge, this is the first end-to-end learned machine-targeted image codec.