Using human gaze in few-shot imitation learning for robot manipulation

Using human gaze in few-shot imitation learning for robot manipulation
复制标题

DOI:
10.1109/iros47612.2022.9981706
复制
发表时间:
2022-10
期刊:
2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Shogo Hamano;Heecheol Kim;Y. Ohmura;Y. Kuniyoshi
Shogo Hamano;Heecheol Kim;Y. Ohmura;Y. Kuniyoshi
中科院分区:
其他
文献类型:
--
作者:
Shogo Hamano;Heecheol Kim;Y. Ohmura;Y. Kuniyoshi

文献摘要

被引文献

相似文献

模仿学习作为一种无需对机器人行为进行编程即可实现复杂机器人控制的方法,受到了广泛的关注。元模仿学习被提出是为了解决模仿学习所面临的数据收集成本高和对新任务的泛化能力低的问题。元模拟通过在训练过程中学习多个任务,可以从少量数据中学习涉及未知对象的新任务。然而,元模仿学习,特别是使用图像,仍然容易受到背景变化的影响,背景占据了输入图像的很大一部分。本研究将人的视线引入基于元模仿学习的机器人控制中。我们创建了一个模型不可知元学习模型,通过在头盔显示器上用眼睛跟踪器测量凝视来预测图像中的凝视位置。使用预测凝视位置周围的图像作为输入使模型对视觉信息的变化具有鲁棒性。通过模拟机器人对任务进行拣选,实验验证了该方法的性能。结果表明,即使在训练和测试过程中对象的颜色或背景图案发生变化,我们提出的方法也比传统的方法具有更强的学习能力,只需从9个演示中学习一个新任务。
Imitation learning has attracted attention as a method for realizing complex robot control without programmed robot behavior. Meta-imitation learning has been proposed to solve the high cost of data collection and low generalizability to new tasks that imitation learning suffers from. Meta-imitation can learn new tasks involving unknown objects from a small amount of data by learning multiple tasks during training. However, meta-imitation learning, especially using images, is still vulnerable to changes in the background, which occupies a large portion of the input image. This study introduces a human gaze into meta-imitation learning-based robot control. We created a model with model-agnostic meta-learning to predict the gaze position from the image by measuring the gaze with an eye tracker in the head-mounted display. Using images around the predicted gaze position as an input makes the model robust to changes in visual information. We experimentally verified the performance of the proposed method through picking tasks using a simulated robot. The results indicate that our proposed method has a greater ability than the conventional method to learn a new task from only 9 demonstrations even if the object's color or the background pattern changes between the training and test.