Video Compressed Sensing Using a Convolutional Neural Network

Video Compressed Sensing Using a Convolutional Neural Network
复制标题

使用卷积神经网络的视频压缩感知

DOI:
10.1109/tcsvt.2020.2978703
复制
发表时间:
2021-02-01
影响因子:
8.4
通讯作者:
Zhao, Debin
Zhao, Debin
中科院分区:
工程技术1区
文献类型:
--
作者:
Shi, Wuzhen;Liu, Shaohui;Zhao, Debin

文献摘要

被引文献

相似文献

近年来,一些基于深度学习的图像压缩感知(CS)方法得到了发展,以较低的计算复杂度实现了较好的重建质量。然而,这些现有的基于深度学习的图像CS方法侧重于探索帧内相关性,而忽略了帧间线索,导致直接应用于视频CS时效率低下。在本文中,我们提出了一种基于卷积神经网络(称为VCSNet)的新型视频CS框架,以探索帧内和帧间的相关性。具体来说,VCSNet将视频序列分成多组图片(GOPs),其中第一帧为关键帧,以比其他非关键帧更高的采样率进行采样。在GOP中,提出了基于卷积层的分块逐帧采样方法,实现了采样矩阵的自动优化。在重建过程中,首先提出了基于线性卷积神经网络的逐帧初始重建,有效地利用了帧内相关性。然后,提出了多层特征补偿的深度重构方法,以多层特征补偿的方式对关键帧进行非关键帧的补偿。这种多层特征补偿允许网络更好地探索帧内和帧间的相关性。在6个基准视频上进行的大量实验表明,VCSNet在客观和主观重建质量上都优于最先进的视频CS方法和基于深度学习的图像CS方法。
Recently, a few image compressed sensing (CS) methods based on deep learning have been developed, which achieve remarkable reconstruction quality with low computational complexity. However, these existing deep learning-based image CS methods focus on exploring intraframe correlation while ignoring interframe cues, resulting in inefficiency when directly applied to video CS. In this paper, we propose a novel video CS framework based on a convolutional neural network (dubbed VCSNet) to explore both intraframe and interframe correlations. Specifically, VCSNet divides the video sequence into multiple groups of pictures (GOPs), of which the first frame is a keyframe that is sampled at a higher sampling ratio than the other nonkeyframes. In a GOP, the block-based framewise sampling by a convolution layer is proposed, which leads to the sampling matrix being automatically optimized. In the reconstruction process, the framewise initial reconstruction by using a linear convolutional neural network is first presented, which effectively utilizes the intraframe correlation. Then, the deep reconstruction with multilevel feature compensation is proposed, which compensates the nonkeyframes with the keyframe in a multilevel feature compensation manner. Such multilevel feature compensation allows the network to better explore both intraframe and interframe correlations. Extensive experiments on six benchmark videos show that VCSNet provides better performance over state-of-the-art video CS methods and deep learning-based image CS methods in both objective and subjective reconstruction quality.