Robust Behavior Cloning with Adversarial Demonstration Detection

Robust Behavior Cloning with Adversarial Demonstration Detection
复制标题

DOI:
10.1109/iros51168.2021.9636203
复制
发表时间:
2021-09
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Mostafa Hussein;Brenda J. Crowe;Madison Clark-Turner;Paul Gesel;Marek Petrik;M. Begum
Mostafa Hussein;Brenda J. Crowe;Madison Clark-Turner;Paul Gesel;Marek Petrik;M. Begum
中科院分区:
其他
文献类型:
--
作者:
Mostafa Hussein;Brenda J. Crowe;Madison Clark-Turner;Paul Gesel;Marek Petrik;M. Begum

文献摘要

相似文献

机器人学中的模仿学习(IL)框架通常假设领域专家的演示总是包含完成任务的正确方法。尽管这一假设在理论上很方便,但在现实世界中,对于IL驱动的机器人来说,它的实用价值有限。现实世界中的专家提供可能包含不正确或潜在不安全的任务方法的演示有很多原因。为了让IL驱动的机器人在现实世界中工作,IL框架需要检测这种对抗性演示,而不是从中学习。本文提出了一个IL框架,它可以自动检测和删除演示集中存在的对抗性演示,因为它直接从专家那里学习任务策略。我们称之为鲁棒最大熵行为克隆(R-MaxEnt)的拟议框架学习了一个将状态映射到动作的随机模型。在这样做的过程中,R-MaxEnt解决了一个最小最大问题,该问题利用模型的熵为不同的演示分配权重,同时为对抗性样本分配较差的权重。我们的实验结果表明,R-MaxEnt在真实和模拟机器人任务中的表现都优于现有的IL方法。
Imitation learning (IL) frameworks in robotics typically assume that a domain expert's demonstration always contains a correct way of doing the task. Despite its theoretical convenience, this assumption has limited practical values for an IL-powered robot in real world. There are many reasons for an expert in the real world to provide demonstrations that may contain incorrect or potentially unsafe way of doing a task. In order for IL-powered robots to work in the real world, IL frameworks need to detect such adversarial demonstrations and not learn from them. This paper proposes an IL framework that can autonomously detect and remove adversarial demonstrations, if they exist in the demonstration set, as it directly learns a task policy from the expert. The proposed framework that we term Robust Maximum Entropy behavior cloning (R-MaxEnt) learns a stochastic model that maps states to actions. In doing so, R-MaxEnt solves a minmax problem that leverages the entropy of the model to assign weights to different demonstrations while assigning poor weights to adversarial samples. Our empirical results show that R-MaxEnt outperforms the existing IL approaches in both real and simulated robotics tasks.