Structured Object-Aware Physics Prediction for Video Modeling and Planning

Structured Object-Aware Physics Prediction for Video Modeling and Planning
复制标题

用于视频建模和规划的结构化对象感知物理预测

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on Learning Representations
影响因子:
--
通讯作者:
K. Kersting
K. Kersting
中科院分区:
--
文献类型:
--
作者:
Jannik Kossen;Karl Stelzner;Marcel Hussing;C. Voelcker;K. Kersting

文献摘要

被引文献

相似文献

当人类观察一个物理系统时,他们可以很容易地定位对象,理解它们的相互作用,并预测未来的行为,即使是在复杂的和以前看不见的相互作用的环境中。然而,对于计算机来说,以无监督的方式从视频中学习这样的模型是一个尚未解决的研究问题。在本文中,我们提出了STOVE,一种新的状态空间模型的视频,明确的原因对象及其位置,速度和相互作用。它通过组合图像模型和动力学模型来构建,并通过重用动力学模型进行推理,加速和正则化训练来改进以前的工作。STOVE在数百个时间步上预测具有令人信服的物理行为的视频,优于以前的无监督模型,甚至接近有监督基线的性能。我们进一步证明了我们的模型作为一个模拟器的样本有效的基于模型的控制在一个任务与大量的相互作用的对象的强度。
When humans observe a physical system, they can easily locate objects, understand their interactions, and anticipate future behavior, even in settings with complicated and previously unseen interactions. For computers, however, learning such models from videos in an unsupervised fashion is an unsolved research problem. In this paper, we present STOVE, a novel state-space model for videos, which explicitly reasons about objects and their positions, velocities, and interactions. It is constructed by combining an image model and a dynamics model in compositional manner and improves on previous work by reusing the dynamics model for inference, accelerating and regularizing training. STOVE predicts videos with convincing physical behavior over hundreds of timesteps, outperforms previous unsupervised models, and even approaches the performance of supervised baselines. We further demonstrate the strength of our model as a simulator for sample efficient model-based control in a task with heavily interacting objects.