Guiding Search in Continuous State-Action Spaces by Learning an Action Sampler From Off-Target Search Experience

Guiding Search in Continuous State-Action Spaces by Learning an Action Sampler From Off-Target Search Experience
复制标题

DOI:
10.1609/aaai.v32i1.12106
复制
发表时间:
2018-04
期刊:
--
影响因子:
--
通讯作者:
Beomjoon Kim;L. Kaelbling;Tomas Lozano-Perez
Beomjoon Kim;L. Kaelbling;Tomas Lozano-Perez
中科院分区:
其他
文献类型:
--
作者:
Beomjoon Kim;L. Kaelbling;Tomas Lozano-Perez

文献摘要

被引文献

相似文献

在机器人技术中,能够在高维连续状态-动作空间中有效地进行长期规划是至关重要的。对于这种复杂的规划问题,无指导的统一采样的行动,直到找到一条路径到一个目标是无可救药的效率低下,和基于梯度的方法往往不符合给定问题的优化流形是不顺利的。在本文中,我们提出了一种方法,通过学习过去的搜索经验的行动采样器,在连续空间中的通用规划者指导搜索。我们使用生成对抗网络(GAN)来表示动作采样器,并解决一个重要问题:搜索经验由相对大量的不在解路径上的动作和相对少量的实际上在解路径上的动作组成。我们引入一种新的技术,基于重要性比估计方法,使用来自非目标分布的样本,使GAN学习更具数据效率。我们提供了理论保证和经验评估在三个具有挑战性的连续机器人规划问题,以说明我们的算法的有效性。
In robotics, it is essential to be able to plan efficiently in high-dimensional continuous state-action spaces for long horizons. For such complex planning problems, unguided uniform sampling of actions until a path to a goal is found is hopelessly inefficient, and gradient-based approaches often fall short when the optimization manifold of a given problem is not smooth. In this paper, we present an approach that guides search in continuous spaces for generic planners by learning an action sampler from past search experience. We use a Generative Adversarial Network (GAN) to represent an action sampler, and address an important issue: search experience consists of a relatively large number of actions that are not on a solution path and a relatively small number of actions that actually are on a solution path. We introduce a new technique, based on an importance-ratio estimation method, for using samples from a non-target distribution to make GAN learning more data-efficient. We provide theoretical guarantees and empirical evaluation in three challenging continuous robot planning problems to illustrate the effectiveness of our algorithm.