Learning Plannable Representations with Causal InfoGAN

Learning Plannable Representations with Causal InfoGAN
复制标题

DOI:
--
复制
发表时间:
2018-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Thanard Kurutach;Aviv Tamar;Ge Yang;Stuart J. Russell;P. Abbeel
Thanard Kurutach;Aviv Tamar;Ge Yang;Stuart J. Russell;P. Abbeel
中科院分区:
其他
文献类型:
--
作者:
Thanard Kurutach;Aviv Tamar;Ge Yang;Stuart J. Russell;P. Abbeel

文献摘要

被引文献

相似文献

近年来,深度生成模型已经被证明可以“想象”令人信服的高维观察,如图像,音频甚至视频,直接从原始数据中学习。在这项工作中,我们问如何想象目标导向的视觉计划-一个合理的观察序列,将动态系统从当前配置转换到期望的目标状态,稍后可以用作控制的参考轨迹。我们专注于具有高维观测的系统,例如图像,并提出了一种自然结合表示学习和规划的方法。我们的框架学习了一个连续观测的生成模型,其中生成过程是由低维规划模型中的过渡和额外的噪声引起的。通过最大化生成的观测值和规划模型中的过渡之间的互信息,我们获得了一个低维表示,最好地解释了数据的因果性质。我们的规划模型的结构,以兼容高效的规划算法,我们提出了几个这样的模型的基础上,无论是离散或连续状态。最后,为了生成可视化计划,我们将当前和目标观测投影到规划模型中的相应状态,规划轨迹,然后使用生成模型将轨迹转换为观测序列。我们展示了我们的方法,想象合理的视觉计划的绳子操作。
In recent years, deep generative models have been shown to 'imagine' convincing high-dimensional observations such as images, audio, and even video, learning directly from raw data. In this work, we ask how to imagine goal-directed visual plans – a plausible sequence of observations that transition a dynamical system from its current configuration to a desired goal state, which can later be used as a reference trajectory for control. We focus on systems with high-dimensional observations, such as images, and propose an approach that naturally combines representation learning and planning. Our framework learns a generative model of sequential observations, where the generative process is induced by a transition in a low-dimensional planning model, and an additional noise. By maximizing the mutual information between the generated observations and the transition in the planning model, we obtain a low-dimensional representation that best explains the causal nature of the data. We structure the planning model to be compatible with efficient planning algorithms, and we propose several such models based on either discrete or continuous states. Finally, to generate a visual plan, we project the current and goal observations onto their respective states in the planning model, plan a trajectory, and then use the generative model to transform the trajectory to a sequence of observations. We demonstrate our method on imagining plausible visual plans of rope manipulation.