Deep multi-scale video prediction beyond mean square error

Deep multi-scale video prediction beyond mean square error
复制标题

DOI:
--
复制
发表时间:
2015-11
期刊:
CoRR
影响因子:
--
通讯作者:
Michaël Mathieu;C. Couprie;Yann LeCun
Michaël Mathieu;C. Couprie;Yann LeCun
中科院分区:
其他
文献类型:
--
作者:
Michaël Mathieu;C. Couprie;Yann LeCun

文献摘要

被引文献

相似文献

学习从视频序列预测未来的图像涉及构建一个内部表示,该内部表示准确地模拟图像的演变,因此在某种程度上,它的内容和动态。这就是为什么像素空间视频预测可能被视为无监督特征学习的一条很有前途的途径。此外,虽然光流一直是计算机视觉中非常研究的问题,但很少有人研究未来的帧预测。尽管如此,许多视觉应用程序可以受益于下一帧视频的知识,这不需要跟踪每一个像素轨迹的复杂性。在这项工作中,我们训练一个卷积网络来生成给定的输入序列的未来帧。针对标准均方误差(MSE)损失函数预测模糊的问题,提出了三种不同的、互补的特征学习策略:多尺度结构、对抗性训练方法和图像梯度差损失函数。我们在UCF101数据集上将我们的预测与基于递归神经网络的不同发表结果进行了比较
Learning to predict future images from a video sequence involves the construction of an internal representation that models the image evolution accurately, and therefore, to some degree, its content and dynamics. This is why pixel-space video prediction may be viewed as a promising avenue for unsupervised feature learning. In addition, while optical flow has been a very studied problem in computer vision for a long time, future frame prediction is rarely approached. Still, many vision applications could benefit from the knowledge of the next frames of videos, that does not require the complexity of tracking every pixel trajectories. In this work, we train a convolutional network to generate future frames given an input sequence. To deal with the inherently blurry predictions obtained from the standard Mean Squared Error (MSE) loss function, we propose three different and complementary feature learning strategies: a multi-scale architecture, an adversarial training method, and an image gradient difference loss function. We compare our predictions to different published results based on recurrent neural networks on the UCF101 dataset