Reinforcement Learning-Based Grasping via One-Shot Affordance Localization and Zero-Shot Contrastive Language-Image Learning

Reinforcement Learning-Based Grasping via One-Shot Affordance Localization and Zero-Shot Contrastive Language-Image Learning
复制标题

DOI:
10.1109/sii58957.2024.10417178
复制
发表时间:
2024-01
期刊:
2024 IEEE/SICE International Symposium on System Integration (SII)
影响因子:
--
通讯作者:
Xiang Long;Luke Beddow;Denis Hadjivelichkov;Andromachi Maria Delfaki;Helge A Wurdemann;D. Kanoulas
Xiang Long;Luke Beddow;Denis Hadjivelichkov;Andromachi Maria Delfaki;Helge A Wurdemann;D. Kanoulas
中科院分区:
其他
文献类型:
--
作者:
Xiang Long;Luke Beddow;Denis Hadjivelichkov;Andromachi Maria Delfaki;Helge A Wurdemann;D. Kanoulas

文献摘要

相似文献

我们提出了一种新的机器人抓取系统,使用笼式夹持器,它结合了一次性的启示定位和零杆物体识别。我们展示了一个集成的系统,需要最少的先验知识,专注于灵活的少数拍摄对象不可知的方法。为了抓住一个新的目标对象,我们使用场景的颜色和深度作为输入,一个类似于目标对象的对象启示的图像,以及一个描述目标对象的最多三个字的文本提示。我们演示了系统使用现实世界的抓YCB基准集的对象,与四个干扰物对象杂乱的场景。总体而言,我们的管道具有96%的可供性定位成功率,62.5%的对象识别和72%的抓取。视频在项目网站上:https://sites.google.com/view/rl-affcorrs-grasp。
We present a novel robotic grasping system using a caging-style gripper, that combines one-shot affordance localization and zero-shot object identification. We demonstrate an integrated system requiring minimal prior knowledge, focusing on flexible few-shot object agnostic approaches. For grasping a novel target object, we use as input the color and depth of the scene, an image of an object affordance similar to the target object, and an up to three-word text prompt describing the target object. We demonstrate the system using real-world grasping of objects from the YCB benchmark set, with four distractor objects cluttering the scene. Overall, our pipeline has a success rate of the affordance localization of 96%, object identification of 62.5%, and grasping of 72%. Videos are on the project website: https://sites.google.com/view/rl-affcorrs-grasp.