A recurrent video quality enhancement framework with multi-granularity frame-fusion and frame difference based attention

A recurrent video quality enhancement framework with multi-granularity frame-fusion and frame difference based attention
复制标题

具有多粒度帧融合和基于帧差异的注意力的循环视频质量增强框架

DOI:
10.1016/j.neucom.2020.12.019
复制
发表时间:
2021-01-04
期刊:
影响因子:
6
通讯作者:
Jiang, Jianmin
Jiang, Jianmin
中科院分区:
计算机科学2区
文献类型:
--
作者:
Huo, Yongkai;Lian, Qiyan;Jiang, Jianmin

文献摘要

被引文献

相似文献

近年来,深度学习在视频恢复方面吸引了大量的研究关注。在现有的方法中,基于单帧的方法在增强目标帧时仅依赖于一个参考帧,而忽略了其余的相邻帧。相比之下,基于多帧的贡献利用滑动窗口中的时间信息,并且现有的递归设计仅采用单个先前增强帧。利用多个原始相邻帧和之前的增强帧来增强视频质量是直观的。在本文中,我们提出了一个循环视频质量增强框架与多粒度帧融合和帧差分基于注意力(REMD)。首先,我们设计了一个基于三维卷积神经网络的编解码器融合模型,该模型对多帧进行多粒度融合。其次,严重的压缩伪影往往出现在压缩帧的边缘和纹理上。我们提出了一种基于帧差的空间注意力方法来增强运动区域的边缘和纹理。最后,一个循环的滑动窗口的设计被认为是利用时间信息在先前的增强帧和随后的相邻帧。实验表明,我们的方法实现了上级性能相比,国家的最先进的贡献,大大降低了空间和计算复杂度。(C)2020爱思唯尔B.V.保留所有权利。
In recent years, deep learning has attracted substantial research attention for video restoration. Among the existing contributions, the single-frame based approaches purely rely on one reference frame and neglect the rest neighboring frames when enhancing a target frame. By contrast, the multi-frame based contributions exploit temporal information in a sliding window and the existing recurrent design only employ a single preceding enhanced frame. It is intuitive to exploit both multiple original neighboring frames and the preceding enhanced frames for video quality enhancement. In this paper, we propose a Recurrent video quality Enhancement framework with Multi-granularity frame-fusion and frame Difference based attention (REMD). Firstly, we devise a three-dimensional convolutional neural network based encoder-decoder fusion model, which fuses multiple frames in multi-granularity. Secondly, severe compression artifacts tend to emerge on the edges and textures of the compressed frames. We propose a frame difference based spatial attention method to intensify the edges and textures of motioning regions. Finally, a recurrent sliding window design is conceived for exploiting the temporal information in preceding enhanced frames and subsequent neighboring frames. Experiments demonstrate that our method achieves superior performance in comparison to the state-of-the-art contributions with substantially reduced spatial and computational complexity. (C) 2020 Elsevier B.V. All rights reserved.