Memory-based gaze prediction in deep imitation learning for robot manipulation

Memory-based gaze prediction in deep imitation learning for robot manipulation
复制标题

DOI:
10.1109/icra46639.2022.9812087
复制
发表时间:
2022-02
期刊:
2022 International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Heecheol Kim;Y. Ohmura;Y. Kuniyoshi
Heecheol Kim;Y. Ohmura;Y. Kuniyoshi
中科院分区:
其他
文献类型:
--
作者:
Heecheol Kim;Y. Ohmura;Y. Kuniyoshi

文献摘要

相似文献

在机器人自主操作中,深度模仿学习是一种很有前途的方法,它不需要硬编码的控制规则。目前,深度模仿学习在机器人操作中的应用仅限于基于当前时间步长状态的反应控制。然而,未来的机器人还需要利用它们在复杂环境中通过经验获得的记忆来解决任务(例如,当机器人被要求在架子上找到以前使用过的物体时)。在这种情况下,简单的深度模仿学习可能会因为复杂环境引起的干扰而失败。我们提出,从顺序视觉输入的凝视预测使机器人能够执行需要记忆的操作任务。该算法采用基于transformer的自注意架构进行注视估计,并基于序列数据实现记忆。通过一个需要记忆之前状态的真实机器人多目标操作任务对该方法进行了评估。
Deep imitation learning is a promising approach that does not require hard-coded control rules in autonomous robot manipulation. The current applications of deep imitation learning to robot manipulation have been limited to reactive control based on the states at the current time step. However, future robots will also be required to solve tasks utilizing their memory obtained by experience in complicated environments (e.g., when the robot is asked to find a previously used object on a shelf). In such a situation, simple deep imitation learning may fail because of distractions caused by complicated environments. We propose that gaze prediction from sequential visual input enables the robot to perform a manipulation task that requires memory. The proposed algorithm uses a Transformer-based self-attention architecture for the gaze estimation based on sequential data to implement memory. The proposed method was evaluated with a real robot multi-object manipulation task that requires memory of the previous states.