Deep Bilateral Learning for Real-Time Image Enhancement

Deep Bilateral Learning for Real-Time Image Enhancement
复制标题

DOI:
10.1145/3072959.3073592
复制
发表时间:
2017-07-01
影响因子:
6.2
通讯作者:
Durand, Fredo
Durand, Fredo
中科院分区:
计算机科学1区
文献类型:
--
作者:
Gharbi, Michael;Chen, Jiawen;Durand, Fredo

文献摘要

被引文献

相似文献

性能是移动图像处理的关键问题。给定一个参考成像管道,甚至是人为调整的图像对,我们寻求重现增强并实现实时评估。为此,我们引入了一种基于双边网格处理和局部仿射颜色变换的神经网络结构。使用成对的输入/输出图像,我们训练卷积神经网络来预测双边空间中局部仿射模型的系数。我们的架构学习做出局部、全局和内容相关的决策,以近似期望的图像转换。在运行时,神经网络消耗输入图像的低分辨率版本,在双边空间中产生一组仿射变换,使用新的切片节点以保持边缘的方式对这些变换进行上采样,然后将这些上采样变换应用于全分辨率图像。我们的算法在几毫秒内处理智能手机上的高分辨率图像,提供1080p分辨率的实时取景器,并在大型图像操作器上匹配最先进的近似技术的质量。与以前的工作不同,我们的模型是离线从数据中训练的,因此不需要在运行时访问原始操作符。这使得我们的模型能够学习复杂的、依赖于场景的转换,而这些转换没有可用的参考实现,比如人类修图师的照片编辑。
Performance is a critical challenge in mobile image processing. Given a reference imaging pipeline, or even human-adjusted pairs of images, we seek to reproduce the enhancements and enable real-time evaluation. For this, we introduce a new neural network architecture inspired by bilateral grid processing and local affine color transforms. Using pairs of input/output images, we train a convolutional neural network to predict the coefficients of a locally-affine model in bilateral space. Our architecture learns to make local, global, and content-dependent decisions to approximate the desired image transformation. At runtime, the neural network consumes a low-resolution version of the input image, produces a set of affine transformations in bilateral space, upsamples those transformations in an edge-preserving fashion using a new slicing node, and then applies those upsampled transformations to the full-resolution image. Our algorithm processes high-resolution images on a smartphone in milliseconds, provides a real-time viewfinder at 1080p resolution, and matches the quality of state-of-the-art approximation techniques on a large class of image operators. Unlike previous work, our model is trained off-line from data and therefore does not require access to the original operator at runtime. This allows our model to learn complex, scene-dependent transformations for which no reference implementation is available, such as the photographic edits of a human retoucher.