An Adversarial Objective for Scalable Exploration

An Adversarial Objective for Scalable Exploration
复制标题

DOI:
10.1109/iros51168.2021.9636298
复制
发表时间:
2020-03
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Bernadette Bucher;Karl Schmeckpeper;N. Matni;Kostas Daniilidis
Bernadette Bucher;Karl Schmeckpeper;N. Matni;Kostas Daniilidis
中科院分区:
其他
文献类型:
--
作者:
Bernadette Bucher;Karl Schmeckpeper;N. Matni;Kostas Daniilidis

文献摘要

相似文献

在许多机器人任务中,收集新的经验是昂贵的,因此确定如何在新的环境中有效地探索,以便在尽可能少的试验中学习尽可能多的知识,是机器人技术的一个重要问题。在本文中,我们提出了一种方法,探索学习动力学模型的目的。我们的主要思想是最小化的评分由一个网络作为一个目标的规划者选择的行动。该预测模型与预测模型联合优化,使我们的主动学习方法能够对观察和动作序列进行采样,从而导致预测模型认为最不现实的预测。如果不对标准机器人平台进行硬件修改以适应其大计算需求,现有的类似探索方法无法在机器人学习中使用的许多预测规划管道中运行,因此我们的对抗探索方法的主要贡献是可扩展性。与领先的基于模型的探索策略相比,我们的对抗性探索方法的性能逐渐提高,因为计算在模拟环境中受到限制。我们进一步证明了我们的对抗方法扩展到机器人操作预测规划管道的能力,在那里我们提高了域转移问题的样本效率和预测性能。
Collecting new experience is costly in many robotic tasks, so determining how to efficiently explore in a new environment to learn as much as possible in as few trials as possible is an important problem for robotics. In this paper, we propose a method for exploring for the purpose of learning a dynamics model. Our key idea is to minimize a score given by a discriminator network as an objective for a planner which chooses actions. This discriminator is optimized jointly with a prediction model and enables our active learning approach to sample sequences of observations and actions which result in predictions considered the least realistic by the discriminator. Comparable existing exploration methods cannot operate in many prediction-planning pipelines used in robotic learning without hardware modifications to standard robotics platforms in order to accommodate their large compute requirements, so the primary contribution of our adversarial exploration method is scalability. We demonstrate progressively increased performance of our adversarial exploration approach compared to leading model-based exploration strategies as compute is restricted in simulated environments. We further demonstrate the ability of our adversarial method to scale to a robotic manipulation prediction-planning pipeline where we improve sample efficiency and prediction performance for a domain transfer problem.