Towards High-Quality and Efficient Video Super-Resolution via Spatial-Temporal Data Overfitting

Towards High-Quality and Efficient Video Super-Resolution via Spatial-Temporal Data Overfitting
复制标题

DOI:
10.1109/cvpr52729.2023.00989
复制
发表时间:
2023-03
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Gen Li;Jie Ji;Minghai Qin;Wei Niu;Bin Ren;F. Afghah;Lin Guo;Xiaolong Ma
Gen Li;Jie Ji;Minghai Qin;Wei Niu;Bin Ren;F. Afghah;Lin Guo;Xiaolong Ma
中科院分区:
其他
文献类型:
--
作者:
Gen Li;Jie Ji;Minghai Qin;Wei Niu;Bin Ren;F. Afghah;Lin Guo;Xiaolong Ma

文献摘要

相似文献

随着深度卷积神经网络(DNN)在计算机视觉的各个领域广泛应用,利用DNN的过拟合能力实现视频分辨率提升已成为现代视频传输系统中的一种新趋势。通过将视频分割成块,并使用超分辨率模型对每个块进行过拟合,服务器在将视频传输给客户端之前对其进行编码,从而实现更好的视频质量和传输效率。然而,为确保良好的过拟合质量,预计需要大量的块,这大大增加了存储量,并在数据传输时消耗更多的带宽资源。另一方面,通过训练优化技术减少块的数量通常需要较高的模型容量,这会显著降低执行速度。为了协调这些问题,我们提出了一种用于高质量且高效的视频分辨率提升任务的新方法,该方法利用时空信息将视频精确地分割成块,从而将块的数量以及模型大小保持在最小值。此外,我们通过一种数据感知联合训练技术将我们的方法改进为一个单一的过拟合模型,这进一步降低了存储需求,且质量下降可忽略不计。我们将我们的模型部署在一款现成的手机上,实验结果表明我们的方法实现了具有高视频质量的实时视频超分辨率。与现有技术相比,我们的方法在实时视频分辨率提升任务中实现了28帧/秒的流速度和41.6的峰值信噪比(PSNR),速度快了14倍,峰值信噪比提高了2.29分贝。代码可在https://github.com/coulsonlee/STDO - CVPR2023.git获取。
As deep convolutional neural networks (DNNs) are widely used in various fields of computer vision, leveraging the overfitting ability of the DNN to achieve video resolution upscaling has become a new trend in the modern video delivery system. By dividing videos into chunks and over-fitting each chunk with a super-resolution model, the server encodes videos before transmitting them to the clients, thus achieving better video quality and transmission efficiency. However, a large number of chunks are expected to ensure good overfitting quality, which substantially increases the storage and consumes more bandwidth resources for data transmission. On the other hand, decreasing the number of chunks through training optimization techniques usually requires high model capacity, which significantly slows down execution speed. To reconcile such, we propose a novel method for high-quality and efficient video resolution upscaling tasks, which leverages the spatial-temporal information to accurately divide video into chunks, thus keeping the number of chunks as well as the model size to minimum. Additionally, we advance our method into a single overfitting model by a data-aware joint training technique. which further reduces the storage requirement with negligible quality drop. We deploy our models on an off-the-shelf mobile phone, and experimental results show that our method achieves real-time video super-resolution with high video quality. Compared with the state-of-the-art, our method achieves 28 fps streaming speed with 41.6 PSNR, which is 14 × faster and 2.29 dB better in the live video resolution upscaling tasks. Code available in https://github.com/coulsonlee/STDO-CVPR2023.git.