SOZIL: Self-Optimal Zero-Shot Imitation Learning

SOZIL: Self-Optimal Zero-Shot Imitation Learning
复制标题

DOI:
10.1109/tcds.2021.3116604
复制
发表时间:
2023-12
影响因子:
5
通讯作者:
Peng Hao;Tao Lu;Shaowei Cui;Junhang Wei;Yinghao Cai;Shuo Wang
Peng Hao;Tao Lu;Shaowei Cui;Junhang Wei;Yinghao Cai;Shuo Wang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Peng Hao;Tao Lu;Shaowei Cui;Junhang Wei;Yinghao Cai;Shuo Wang

文献摘要

相似文献

零样本模仿学习已经证明了其在人类参与较少的情况下学习复杂机器人任务的优越性。最近的研究表明,在机器人严格遵循学习的逆模型演示的情况下,其性能令人信服。然而,当演示不理想时,这些方法很难在模仿中获得令人满意的性能,并且学习到的逆模型的学习容易受到标签模糊问题的影响。在本文中,我们提出自优化零样本模仿学习(SOZIL)来解决这些问题。 SOZIL 的贡献是双重的。首先,目标一致性损失(GCL)旨在从探索数据中学习多步骤目标条件策略。 GCL通过直接使用目标状态作为监督,解决了轨迹和动作多样性引起的标签模糊问题。其次,开发基于估计的关键帧提取(EKE)来优化演示。我们将关键帧提取过程表述为次优控制下的路径优化问题。通过预测学习策略在执行任意两个状态转换时的性能,EKE 创建一个包含所有候选路径的有向图,并通过解决该图的最短路径问题来提取关键帧。此外,还通过各种模拟和真实机器人操作实验(例如线束组装、绳索操作和块移动)对所提出的方法进行了评估。实验结果表明,SOZIL 实现了比基线更高的成功率和操作效率。
Zero-shot imitation learning has demonstrated its superiority to learn complex robotic tasks with less human participation. Recent studies show convincing performance under the condition that the robot follows the demonstration strictly by the learned inverse model. However, these methods are difficult to achieve satisfactory performance in imitation when the demonstration is suboptimal, and the learning of the learned inverse models is vulnerable to label ambiguity issues. In this article, we propose self-optimal zero-shot imitation learning (SOZIL) to tackle these problems. The contribution of SOZIL is twofold. First, goal consistency loss (GCL) is designed to learn the multistep goal-conditioned policy from exploration data. By directly using the goal state as supervision, GCL solves the label ambiguity problem caused by trajectory and action diversity. Second, estimation-based keyframe extraction (EKE) is developed to optimize demonstrations. We formulate the keyframe extraction process as a path optimization problem under suboptimal control. By predicting the performance of the learned policy in executing transitions of any two states, EKE creates a directed graph containing all candidate paths and extracts keyframes by solving the graph’s shortest path problem. Furthermore, the proposed method is evaluated with various simulated and real-world robotic manipulating experiments, such as cable harness assembly, rope manipulation, and block moving. Experimental results show that SOZIL achieves a superior success rate and manipulation efficiency than baselines.