Universal Planning Networks

Universal Planning Networks
复制标题

DOI:
--
复制
发表时间:
2018-04
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Srinivas;A. Jabri;P. Abbeel;S. Levine;Chelsea Finn
A. Srinivas;A. Jabri;P. Abbeel;S. Levine;Chelsea Finn
中科院分区:
其他
文献类型:
--
作者:
A. Srinivas;A. Jabri;P. Abbeel;S. Levine;Chelsea Finn

文献摘要

被引文献

相似文献

复杂视觉运动控制的一个关键挑战是学习对指定目标、规划和泛化有效的抽象表示。为此,我们引入通用规划网络(UPN)。 UPN 将可区分的规划嵌入到目标导向的政策中。该规划计算在潜在空间中展开前向模型,并通过梯度下降轨迹优化推断出最佳行动计划。梯度下降计划过程及其底层表示是端到端学习的,以直接优化监督模仿学习目标。我们发现,学习到的表示不仅对通过基于梯度的轨迹优化进行目标导向的视觉模仿有效,而且还可以提供使用图像指定目标的度量。可以利用学习到的表示来指定基于距离的奖励,以达到无模型强化学习的新目标状态,从而在解决通过基于图像的目标描述的新任务时实现更有效的学习。我们能够在形态和驱动能力显着不同的机器人之间成功转移视觉运动规划策略。
A key challenge in complex visuomotor control is learning abstract representations that are effective for specifying goals, planning, and generalization. To this end, we introduce universal planning networks (UPN). UPNs embed differentiable planning within a goal-directed policy. This planning computation unrolls a forward model in a latent space and infers an optimal action plan through gradient descent trajectory optimization. The plan-by-gradient-descent process and its underlying representations are learned end-to-end to directly optimize a supervised imitation learning objective. We find that the representations learned are not only effective for goal-directed visual imitation via gradient-based trajectory optimization, but can also provide a metric for specifying goals using images. The learned representations can be leveraged to specify distance-based rewards to reach new target states for model-free reinforcement learning, resulting in substantially more effective learning when solving new tasks described via image-based goals. We were able to achieve successful transfer of visuomotor planning strategies across robots with significantly different morphologies and actuation capabilities.