DART: Noise Injection for Robust Imitation Learning

DART: Noise Injection for Robust Imitation Learning
复制标题

DOI:
--
复制
发表时间:
2017-03
期刊:
--
影响因子:
--
通讯作者:
Michael Laskey;Jonathan Lee;Roy Fox;A. Dragan;Ken Goldberg
Michael Laskey;Jonathan Lee;Roy Fox;A. Dragan;Ken Goldberg
中科院分区:
其他
文献类型:
--
作者:
Michael Laskey;Jonathan Lee;Roy Fox;A. Dragan;Ken Goldberg

文献摘要

相似文献

模仿学习的一种方法是行为克隆,其中机器人观察主管并推断控制策略。这种“偏离策略”方法的一个已知问题是,当机器人偏离主管的演示时,其错误会加剧。在政策上,技术通过迭代地收集当前机器人政策的纠正措施来缓解这一问题。然而,这些技术对于人类监管者来说可能很乏味,增加了大量的计算负担,并且可能在训练期间访问危险状态。我们提出了一种离策略方法,在演示时将噪音注入主管的策略中。这迫使主管演示如何从错误中恢复。我们提出了一种新算法,DART(增强机器人轨迹的干扰),该算法收集注入噪声的演示,并优化噪声水平以近似数据收集期间机器人训练策略的误差。我们在两个领域将 DART 与 DAgger 和行为克隆进行比较:在 MuJoCo 任务(Walker、Humanoid、Hopper、Half-Cheetah)上使用算法主管进行模拟,以及在物理实验中使用人类主管训练丰田 HSR 机器人在杂乱中执行抓取操作。对于像 Humanoid 这样的高维任务,DART 的计算时间最多可以快 $3x$,并且在训练期间仅将监督者的累积奖励减少 $5\%$,而 DAgger 执行的策略的累积奖励比监督者少 $80\%$。在杂乱任务的抓取中,DART 比行为克隆平均获得了 62\%$ 的性能提升。
One approach to Imitation Learning is Behavior Cloning, in which a robot observes a supervisor and infers a control policy. A known problem with this "off-policy" approach is that the robot's errors compound when drifting away from the supervisor's demonstrations. On-policy, techniques alleviate this by iteratively collecting corrective actions for the current robot policy. However, these techniques can be tedious for human supervisors, add significant computation burden, and may visit dangerous states during training. We propose an off-policy approach that injects noise into the supervisor's policy while demonstrating. This forces the supervisor to demonstrate how to recover from errors. We propose a new algorithm, DART (Disturbances for Augmenting Robot Trajectories), that collects demonstrations with injected noise, and optimizes the noise level to approximate the error of the robot's trained policy during data collection. We compare DART with DAgger and Behavior Cloning in two domains: in simulation with an algorithmic supervisor on the MuJoCo tasks (Walker, Humanoid, Hopper, Half-Cheetah) and in physical experiments with human supervisors training a Toyota HSR robot to perform grasping in clutter. For high dimensional tasks like Humanoid, DART can be up to $3x$ faster in computation time and only decreases the supervisor's cumulative reward by $5\%$ during training, whereas DAgger executes policies that have $80\%$ less cumulative reward than the supervisor. On the grasping in clutter task, DART obtains on average a $62\%$ performance increase over Behavior Cloning.