Residual Policy Learning for Shared Autonomy

Residual Policy Learning for Shared Autonomy
复制标题

DOI:
10.15607/rss.2020.xvi.072
复制
发表时间:
2020-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Charles B. Schaff;Matthew R. Walter
Charles B. Schaff;Matthew R. Walter
中科院分区:
其他
文献类型:
--
作者:
Charles B. Schaff;Matthew R. Walter

文献摘要

被引文献

相似文献

共享自主为人-机器人协作提供了一个有效的框架,利用人类和机器人的互补优势来实现共同的目标。许多现有的共享自主方法都做出了限制性的假设,即目标空间、环境动态或人类策略是先验已知的,或者仅限于离散的动作空间,从而阻止了这些方法扩展到复杂的现实世界环境。我们提出了一种用于共享自治的无模型、剩余策略学习算法,它减少了对这些假设的需要。我们的代理经过培训,以最小限度地调整人类的行为,从而满足一组与目标无关的约束。我们在两个连续的控制环境中测试了我们的方法:月球着陆器,一个2D飞行控制域,和一个六自由度四旋翼到达任务。在与人类和代理飞行员的实验中,我们的方法显著提高了任务性能,而不需要知道人类的目标超出了限制。这些结果突出了无模型深度强化学习能够在几乎不了解用户意图的情况下实现适合于连续控制设置的辅助代理的能力。
Shared autonomy provides an effective framework for human-robot collaboration that takes advantage of the complementary strengths of humans and robots to achieve common goals. Many existing approaches to shared autonomy make restrictive assumptions that the goal space, environment dynamics, or human policy are known a priori, or are limited to discrete action spaces, preventing those methods from scaling to complicated real world environments. We propose a model-free, residual policy learning algorithm for shared autonomy that alleviates the need for these assumptions. Our agents are trained to minimally adjust the human's actions such that a set of goal-agnostic constraints are satisfied. We test our method in two continuous control environments: Lunar Lander, a 2D flight control domain, and a 6-DOF quadrotor reaching task. In experiments with human and surrogate pilots, our method significantly improves task performance without any knowledge of the human's goal beyond the constraints. These results highlight the ability of model-free deep reinforcement learning to realize assistive agents suited to continuous control settings with little knowledge of user intent.