Zooming Slow-Mo: Fast and Accurate One-Stage Space-Time Video Super-Resolution

Zooming Slow-Mo: Fast and Accurate One-Stage Space-Time Video Super-Resolution
复制标题

DOI:
10.1109/cvpr42600.2020.00343
复制
发表时间:
2020-02
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Xiaoyu Xiang;Yapeng Tian;Yulun Zhang;Y. Fu;J. Allebach;Chenliang Xu
Xiaoyu Xiang;Yapeng Tian;Yulun Zhang;Y. Fu;J. Allebach;Chenliang Xu
中科院分区:
其他
文献类型:
--
作者:
Xiaoyu Xiang;Yapeng Tian;Yulun Zhang;Y. Fu;J. Allebach;Chenliang Xu

文献摘要

被引文献

相似文献

在本文中,我们探讨了时空视频超分辨率任务,其目的是从低帧率(LFR),低分辨率(LR)视频生成高分辨率(HR)慢动作视频。一个简单的解决方案是将其分为两个子任务:视频帧插值(VFI)和视频超分辨率(VSR)。然而,时间内插和空间超分辨率是内在相关的这项任务。两阶段方法不能充分利用自然属性。此外,现有技术的VFI或VSR网络需要大的帧合成或重构模块来预测高质量的视频帧,这使得两阶段方法具有大的模型尺寸,因此是耗时的。为了克服这些问题,我们提出了一个单级时空视频超分辨率框架,它直接从LFR,LR视频合成HR慢动作视频。而不是合成丢失的LR视频帧的VFI网络做,我们首先在时间上插值LR帧的功能,在丢失的LR视频帧捕获本地的时间上下文所提出的特征时间插值网络。然后,我们提出了一个可变形的ConvLSTM对齐和聚合的时间信息,同时更好地利用全局时间上下文。最后,采用深度重构网络对HR慢动作视频帧进行预测。在基准数据集上进行的大量实验表明,所提出的方法不仅实现了更好的定量和定性性能,而且比最近的两阶段最先进方法(例如,DAIN+EDVR和DAIN+RBPN。
In this paper, we explore the space-time video super-resolution task, which aims to generate a high-resolution (HR) slow-motion video from a low frame rate (LFR), low-resolution (LR) video. A simple solution is to split it into two sub-tasks: video frame interpolation (VFI) and video super-resolution (VSR). However, temporal interpolation and spatial super-resolution are intra-related in this task. Two-stage methods cannot fully take advantage of the natural property. In addition, state-of-the-art VFI or VSR networks require a large frame-synthesis or reconstruction module for predicting high-quality video frames, which makes the two-stage methods have large model sizes and thus be time-consuming. To overcome the problems, we propose a one-stage space-time video super-resolution framework, which directly synthesizes an HR slow-motion video from an LFR, LR video. Rather than synthesizing missing LR video frames as VFI networks do, we firstly temporally interpolate LR frame features in missing LR video frames capturing local temporal contexts by the proposed feature temporal interpolation network. Then, we propose a deformable ConvLSTM to align and aggregate temporal information simultaneously for better leveraging global temporal contexts. Finally, a deep reconstruction network is adopted to predict HR slow-motion video frames. Extensive experiments on benchmark datasets demonstrate that the proposed method not only achieves better quantitative and qualitative performance but also is more than three times faster than recent two-stage state-of-the-art methods, e.g., DAIN+EDVR and DAIN+RBPN.