FitVid: Overfitting in Pixel-Level Video Prediction

FitVid: Overfitting in Pixel-Level Video Prediction
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
M. Babaeizadeh;M. Saffar;Suraj Nair;S. Levine;Chelsea Finn;D. Erhan
M. Babaeizadeh;M. Saffar;Suraj Nair;S. Levine;Chelsea Finn;D. Erhan
中科院分区:
其他
文献类型:
--
作者:
M. Babaeizadeh;M. Saffar;Suraj Nair;S. Levine;Chelsea Finn;D. Erhan

文献摘要

被引文献

相似文献

一种能够预测接下来会发生什么的智能体,无需额外训练,通过规划就能执行各种任务。此外,这样的智能体能够在内部表征现实世界的复杂动态,因此能够获得对各种视觉感知任务有用的表征。这使得根据观察到的过去以及可能的未来动作来预测视频的未来帧成为一项有趣的任务,尽管最近取得了许多进展,但这项任务仍然极具挑战性。现有的视频预测模型在简单的狭义基准测试上显示出有希望的结果,但它们在具有更复杂动态或更广泛领域的现实生活数据集上生成的预测质量较低。越来越多的证据表明,对训练数据的欠拟合是预测质量低的主要原因之一。在本文中,我们认为当前视频模型中参数的低效使用是欠拟合的主要原因。因此,我们引入了一种名为FitVid的新架构,它能够在常见基准测试上严重过拟合,同时参数数量与当前最先进的模型相似。我们分析了过拟合的后果,说明了它如何能产生意想不到的结果,比如通过重复训练数据生成高质量的输出,以及如何使用现有的图像增强技术来缓解它。结果,FitVid在四个不同的视频预测基准测试的四个不同指标上优于当前最先进的模型。
An agent that is capable of predicting what happens next can perform a variety of tasks through planning with no additional training. Furthermore, such an agent can internally represent the complex dynamics of the real-world and therefore can acquire a representation useful for a variety of visual perception tasks. This makes predicting the future frames of a video, conditioned on the observed past and potentially future actions, an interesting task which remains exceptionally challenging despite many recent advances. Existing video prediction models have shown promising results on simple narrow benchmarks but they generate low quality predictions on real-life datasets with more complicated dynamics or broader domain. There is a growing body of evidence that underfitting on the training data is one of the primary causes for the low quality predictions. In this paper, we argue that the inefficient use of parameters in the current video models is the main reason for underfitting. Therefore, we introduce a new architecture, named FitVid, which is capable of severe overfitting on the common benchmarks while having similar parameter count as the current state-of-the-art models. We analyze the consequences of overfitting, illustrating how it can produce unexpected outcomes such as generating high quality output by repeating the training data, and how it can be mitigated using existing image augmentation techniques. As a result, FitVid outperforms the current state-of-the-art models across four different video prediction benchmarks on four different metrics.