One-shot Visual Imitation via Attributed Waypoints and Demonstration Augmentation

One-shot Visual Imitation via Attributed Waypoints and Demonstration Augmentation
复制标题

DOI:
10.1109/icra48891.2023.10160944
复制
发表时间:
2023-02
期刊:
2023 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Matthew Chang;Saurabh Gupta
Matthew Chang;Saurabh Gupta
中科院分区:
其他
文献类型:
--
作者:
Matthew Chang;Saurabh Gupta

文献摘要

相似文献

在本文中,我们分析了现有技术的行为,并设计了新的解决方案,一次性的视觉模仿的问题。在这种情况下,智能体必须解决一个新任务的一个新实例,只需一个视觉演示。我们的分析表明,当前的方法由于三个错误而达不到要求:纯粹离线训练产生的Dagger问题、与对象交互时的最后一厘米错误,以及与任务上下文而不是实际任务的不匹配。这激发了我们的模块化方法的设计,其中我们a)将任务推理(做什么)与任务执行(如何做)分开,以及B)开发数据增强和生成技术以减轻失配。前者允许我们利用手工制作的运动原语来执行任务,从而避免了Dagger问题和最后的厘米误差,而后者则使模型专注于任务而不是任务上下文。我们的模型在最近的两个基准测试中获得了100%和48%的成功率,分别比目前的最先进水平提高了90%和20%。
In this paper, we analyze the behavior of existing techniques and design new solutions for the problem of one-shot visual imitation. In this setting, an agent must solve a novel instance of a novel task given just a single visual demonstration. Our analysis reveals that current methods fall short because of three errors: the DAgger problem arising from purely offline training, last centimeter errors in interacting with objects, and mis-fitting to the task context rather than to the actual task. This motivates the design of our modular approach where we a) separate out task inference (what to do) from task execution (how to do it), and b) develop data augmentation and generation techniques to mitigate mis-fitting. The former allows us to leverage hand-crafted motor primitives for task execution which side-steps the DAgger problem and last centimeter errors, while the latter gets the model to focus on the task rather than the task context. Our model gets 100% and 48% success rates on two recent benchmarks, improving upon the current state-of-the-art by absolute 90% and 20% respectively.