MVStylizer: an efficient edge-assisted video photorealistic style transfer system for mobile phones

MVStylizer: an efficient edge-assisted video photorealistic style transfer system for mobile phones
复制标题

DOI:
10.1145/3397166.3409140
复制
发表时间:
2020-05
期刊:
Proceedings of the Twenty-First International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing
影响因子:
--
通讯作者:
Ang Li;Chunpeng Wu;Yiran Chen;Bin Ni
Ang Li;Chunpeng Wu;Yiran Chen;Bin Ni
中科院分区:
其他
文献类型:
--
作者:
Ang Li;Chunpeng Wu;Yiran Chen;Bin Ni

文献摘要

相似文献

最近的研究在实现图像的神经风格转移方面取得了很大进展,这意味着将图像转换为期望的风格。许多用户开始使用他们的移动的手机来记录他们的日常生活,然后编辑并与其他用户分享所捕获的图像和视频。然而,直接将现有的风格转移方法应用于视频,即,逐帧传送视频的风格需要非常大量的计算资源。在移动的手机上进行视频的风格转移在技术上仍然是负担不起的。为了解决这一挑战,我们提出MVStylizer,一个有效的边缘辅助逼真的视频风格传输系统的移动的手机。与逐帧执行风格化不同,只有原始视频中的关键帧由边缘服务器上预先训练的深度神经网络(DNN)处理,而其余的风格化中间帧则由我们设计的基于光流的帧插值算法在移动的手机上生成。还提出了一个元平滑模块,同时将风格化帧放大到任意分辨率,并去除这些放大帧中与风格转移相关的失真。此外,为了不断提高DNN模型在边缘服务器上的性能,我们采用联合学习方案,利用从移动的客户端收集的数据不断重新训练边缘服务器上的每个DNN模型,并与云服务器上的全局DNN模型同步。这种方案有效地利用了从各种移动的客户端收集的数据的多样性,并有效地提高了系统性能。我们的实验表明,与最先进的方法相比,MVStylizer可以生成具有更好视觉质量的风格化视频,同时对1920×1080视频实现75.5倍的加速比。
Recent research has made great progress in realizing neural style transfer of images, which denotes transforming an image to a desired style. Many users start to use their mobile phones to record their daily life, and then edit and share the captured images and videos with other users. However, directly applying existing style transfer approaches on videos, i.e., transferring the style of a video frame by frame, requires an extremely large amount of computation resources. It is still technically unaffordable to perform style transfer of videos on mobile phones. To address this challenge, we propose MVStylizer, an efficient edge-assisted photorealistic video style transfer system for mobile phones. Instead of performing stylization frame by frame, only key frames in the original video are processed by a pre-trained deep neural network (DNN) on edge servers, while the rest of stylized intermediate frames are generated by our designed optical-flow-based frame interpolation algorithm on mobile phones. A meta-smoothing module is also proposed to simultaneously upscale a stylized frame to arbitrary resolution and remove style transfer related distortions in these upscaled frames. In addition, for the sake of continuously enhancing the performance of the DNN model on the edge server, we adopt a federated learning scheme to keep retraining each DNN model on the edge server with collected data from mobile clients and syncing with a global DNN model on the cloud server. Such a scheme effectively leverages the diversity of collected data from various mobile clients and efficiently improves the system performance. Our experiments demonstrate that MVStylizer can generate stylized videos with an even better visual quality compared to the state-of-the-art method while achieving 75.5× speedup for 1920×1080 videos.