Fixture-Aware DDQN for Generalized Environment-Enabled Grasping

Fixture-Aware DDQN for Generalized Environment-Enabled Grasping
复制标题

DOI:
10.1109/iros47612.2022.9982182
复制
发表时间:
2022-10
期刊:
2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Eddie Sasagawa;Changhyun Choi
Eddie Sasagawa;Changhyun Choi
中科院分区:
其他
文献类型:
--
作者:
Eddie Sasagawa;Changhyun Choi

文献摘要

相似文献

本文扩展了当夹具(例如,墙,重的物体)被利用。之前解决这个问题的工作是有限的,因为所使用的网络隐式地学习特定的目标和固定装置来利用。然而,可用夹具的概念在不同的环境中可能会有所不同,有时没有任何明显的差异。在本文中,我们提出了一种方法来放宽这一限制,并进一步处理环境中的夹具位置是未知的。这个问题被制定为视觉启示学习在一个部分可观察的设置。我们提出了一种自监督强化学习算法,夹具感知双深度Q网络(FA-DDQN),它处理场景观察,以1)基于参考图像识别目标对象,2)基于与环境的交互区分可能的夹具,最后3)融合信息以生成视觉启示图,以指导机器人成功地进行Slide-to-Wall抓取。我们在模拟和真实的机器人实验中演示了我们提出的解决方案,以表明除了比基线取得更高的成功外,它还对具有不可见对象配置的新场景执行零次泛化。
This paper expands on the problem of grasping an object that can only be grasped by a single parallel gripper when a fixture (e.g., wall, heavy object) is harnessed. Preceding work that tackle this problem are limited in that the employed networks implicitly learn specific targets and fixtures to leverage. However, the notion of a usable fixture can vary in different environments, at times without any outwardly noticeable differences. In this paper, we propose a method to relax this limitation and further handle environments where the fixture location is unknown. The problem is formulated as visual affordance learning in a partially observable setting. We present a self-supervised reinforcement learning algorithm, Fixture-Aware Double Deep Q-Network (FA-DDQN), that processes the scene observation to 1) identify the target object based on a reference image, 2) distinguish possible fixtures based on interaction with the environment, and finally 3) fuse the information to generate a visual affordance map to guide the robot to successful Slide-to-Wall grasps. We demonstrate our proposed solution in simulation and in real robot experiments to show that in addition to achieving higher success than baselines, it also performs zero-shot generalization to novel scenes with unseen object configurations.