Learning from Interventions: Human-robot interaction as both explicit and implicit feedback

Learning from Interventions: Human-robot interaction as both explicit and implicit feedback
复制标题

DOI:
10.15607/rss.2020.xvi.055
复制
发表时间:
2020-07
期刊:
Robotics: Science and Systems XVI
影响因子:
--
通讯作者:
Jonathan Spencer;Sanjiban Choudhury;Matt Barnes;Matt Schmittle;M. Chiang;P. Ramadge;S. Srinivasa
Jonathan Spencer;Sanjiban Choudhury;Matt Barnes;Matt Schmittle;M. Chiang;P. Ramadge;S. Srinivasa
中科院分区:
其他
文献类型:
--
作者:
Jonathan Spencer;Sanjiban Choudhury;Matt Barnes;Matt Schmittle;M. Chiang;P. Ramadge;S. Srinivasa

文献摘要

被引文献

相似文献

如果机器人要在现实世界中完成众多任务,从无缝的人机交互中进行可扩展的机器人学习至关重要。当前的模仿学习方法存在两个缺陷之一。一方面,它们仅仅依赖离线策略的人类演示,这在某些情况下会导致训练 - 测试分布不匹配。另一方面,它们让人类给学习者所经历的每个状态都进行标注,这在很多应用中是不切实际的。我们认为从专家干预中进行交互学习兼具两者的优点。我们的关键见解是,任何数量的专家反馈,无论是通过干预还是不干预,都提供了有关当前状态质量、动作最优性或两者的信息。我们将其形式化为对学习者价值函数的一种约束,并且我们可以使用无悔在线学习技术有效地学习它。我们将我们的方法称为专家干预学习(EIL),并在一个由人类专家参与的真实和模拟驾驶任务中对其进行评估,在该任务中,它仅通过几百个(约一分钟)专家控制的样本就从头开始学习避障。
—Scalable robot learning from seamless human-robot interaction is critical if robots are to solve a multitude of tasks in the real world. Current approaches to imitation learning suffer from one of two drawbacks. On the one hand, they rely solely on off-policy human demonstration, which in some cases leads to a mismatch in train-test distribution. On the other, they burden the human to label every state the learner visits, rendering it impractical in many applications. We argue that learning interactively from expert interventions enjoys the best of both worlds. Our key insight is that any amount of expert feedback, whether by intervention or non-intervention, provides information about the quality of the current state, the optimality of the action, or both. We formalize this as a constraint on the learner’s value function, which we can efficiently learn using no regret, online learning techniques. We call our approach Expert Intervention Learning (EIL), and evaluate it on a real and simulated driving task with a human expert, where it learns collision avoidance from scratch with just a few hundred samples (about one minute) of expert control.