SCALOR: Generative World Models with Scalable Object Representations

SCALOR: Generative World Models with Scalable Object Representations
复制标题

SCALOR:具有可扩展对象表示的生成世界模型

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on Learning Representations
影响因子:
--
通讯作者:
Sungjin Ahn
Sungjin Ahn
中科院分区:
--
文献类型:
--
作者:
Jindong Jiang;Sepehr Janghorbani;Gerard de Melo;Sungjin Ahn

文献摘要

被引文献

相似文献

场景中对象密度的可扩展性是无监督顺序面向对象表示学习的主要挑战。大多数以前的模型已经被证明只能在有几个对象的场景中工作。在本文中,我们提出了SCALOR,一个概率生成世界模型学习可缩放的面向对象的表示的视频。与以前的最先进的模型相比,SCALOR提出的空间并行注意和建议拒绝机制,可以处理数量级更大的对象。此外,我们引入了一个背景模块,允许SCALOR模型复杂的动态背景以及场景中的许多前景对象。我们证明,SCALOR可以处理拥挤的场景,包含多达一百个对象,同时联合建模复杂的动态背景。重要的是,SCALOR是第一个无监督对象表示模型,可用于包含数十个移动对象的自然场景。
Scalability in terms of object density in a scene is a primary challenge in unsupervised sequential object-oriented representation learning. Most of the previous models have been shown to work only on scenes with a few objects. In this paper, we propose SCALOR, a probabilistic generative world model for learning SCALable Object-oriented Representation of a video. With the proposed spatially parallel attention and proposal-rejection mechanisms, SCALOR can deal with orders of magnitude larger numbers of objects compared to the previous state-of-the-art models. Additionally, we introduce a background module that allows SCALOR to model complex dynamic backgrounds as well as many foreground objects in the scene. We demonstrate that SCALOR can deal with crowded scenes containing up to a hundred objects while jointly modeling complex dynamic backgrounds. Importantly, SCALOR is the first unsupervised object representation model shown to work for natural scenes containing several tens of moving objects.