Deep Learning-Based Perceptual Video Quality Enhancement for 3D Synthesized View

Deep Learning-Based Perceptual Video Quality Enhancement for 3D Synthesized View
复制标题

基于深度学习的 3D 合成视图感知视频质量增强

DOI:
10.1109/tcsvt.2022.3147788
复制
发表时间:
2022-08
影响因子:
8.4
通讯作者:
Weisi Lin
Weisi Lin
中科院分区:
工程技术1区
文献类型:
--
作者:
Huan Zhang;Yun Zhang;Linwei Zhu;Weisi Lin

文献摘要

相似文献

由于深度视频中视图之间的遮挡和时间不一致,基于深度图像渲染的 3D 合成视频会出现时空失真。在本文中,我们提出了一种基于深度卷积神经网络(CNN)的合成视频去噪算法,以减少时间闪烁失真并提高 3D 合成视频的感知质量。首先,我们分析时空失真,并将消除时空失真的模型作为感知视频去噪问题。然后,提出了一种基于深度学习的合成视频去噪网络,其中从合成视频质量度量导出CNN友好的时空损失函数,并与单个图像去噪网络架构集成。最后,基于现有的基于CNN的去噪模型,开发了特定方案,即特定的合成视频去噪网络(SynVD-Nets)和通用方案,即通用SynVD-Net(GSynVD-Net),以更有效地处理具有不同失真级别的合成视频。实验结果表明,所提出的 SynVD-Net 和 GSynVD-Net 可以优于基于深度学习的同类方法和传统的去噪方法,并显着提高 3D 合成视频的感知质量。
Due to occlusion among views and temporal inconsistency in depth video, spatio-temporal distortion occurs in 3D synthesized video with depth image-based rendering. In this paper, we propose a deep Convolutional Neural Network (CNN)-based synthesized video denoising algorithm to reduce temporal flicker distortion and improve perceptual quality of 3D synthesized video. First, we analyze the spatio-temporal distortion, and model eliminating spatio-temporal distortion as a perceptual video denoising problem. Then, a deep learning-based synthesized video denoising network is proposed, in which a CNN-friendly spatio-temporal loss function is derived from a synthesized video quality metric and integrated with a single image denoising network architecture. Finally, specific schemes, i.e., specific Synthesized Video Denoising Networks (SynVD-Nets), and a general scheme, i.e., General SynVD-Net (GSynVD-Net), based on existing CNN-based denoising models, are developed to handle synthesized video with different distortion levels more effectively. Experimental results show that the proposed SynVD-Net and GSynVD-Net can outperform deep learning-based counterparts and conventional denoising methods, and significantly enhance perceptual quality of 3D synthesized video.