Goal-aware generative adversarial imitation learning from imperfect demonstration for robotic cloth manipulation

Goal-aware generative adversarial imitation learning from imperfect demonstration for robotic cloth manipulation
复制标题

DOI:
10.1016/j.robot.2022.104264
复制
发表时间:
2022-09-20
影响因子:
4.3
通讯作者:
Matsubara, Takamitsu
Matsubara, Takamitsu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Tsurumine, Yoshihisa;Matsubara, Takamitsu

文献摘要

被引文献

相似文献

生成性对抗模仿学习(GAIL)可以学习策略,而不需要明确定义来自演示的奖励函数。Gail有可能通过高维观察作为输入来学习政策,例如图像。通过将Gail应用到真正的机器人上,也许可以获得洗衣服、叠衣服、做饭和清洁等日常活动的机器人策略。然而,由于错误,人类演示数据往往是不完美的,这降低了所产生的政策的性能。我们通过关注以下特征来解决这个问题:(1)许多机器人任务是到达目标的任务,(2)在演示数据中标记这样的目标状态相对容易。考虑到这些,本文提出了目标感知的生成性对抗模仿学习(GA-GAIL),它通过引入第二个鉴别器来训练策略,以与指示示范数据的第一个鉴别器并行地区分目标状态。这扩展了一个标准的Gail框架,通过一个促进实现目标状态的目标状态鉴别器,更有力地学习理想的政策,即使是从不完美的演示中也是如此。此外,GA-Gail算法还使用了最大熵深度P-网络(EDPN)作为生成器,在策略更新中同时考虑了光滑性和因果性,从而实现了从两个鉴别器进行稳定的策略学习。我们提出的方法被成功地应用于两个真实的机器人布料操纵任务:翻转手帕和折叠衣服。我们证实,它学习布料操纵策略,而不是任务特定的奖励函数设计。实际实验的视频可在此URL上找到。(C)2022爱思唯尔B.V.保留所有权利。
Generative Adversarial Imitation Learning (GAIL) can learn policies without explicitly defining the reward function from demonstrations. GAIL has the potential to learn policies with high-dimensional observations as input, e.g., images. By applying GAIL to a real robot, perhaps robot policies can be obtained for daily activities like washing, folding clothes, cooking, and cleaning. However, human demonstration data are often imperfect due to mistakes, which degrade the performance of the resulting policies. We address this issue by focusing on the following features: (1) many robotic tasks are goal-reaching tasks, and (2) labeling such goal states in demonstration data is relatively easy. With these in mind, this paper proposes Goal-Aware Generative Adversarial Imitation Learning (GA-GAIL), which trains a policy by introducing a second discriminator to distinguish the goal state in parallel with the first discriminator that indicates the demonstration data. This extends a standard GAIL framework to more robustly learn desirable policies even from imperfect demonstrations through a goal-state discriminator that promotes achieving the goal state. Furthermore, GA-GAIL employs the Entropy -maximizing Deep P-Network (EDPN) as a generator, which considers both the smoothness and causal entropy in the policy update, to achieve stable policy learning from two discriminators. Our proposed method was successfully applied to two real-robotic cloth-manipulation tasks: turning a handkerchief over and folding clothes. We confirmed that it learns cloth-manipulation policies without task-specific reward function design. Video of the real experiments are available at this URL.(c) 2022 Elsevier B.V. All rights reserved.