KINet: Unsupervised Forward Models for Robotic Pushing Manipulation

KINet: Unsupervised Forward Models for Robotic Pushing Manipulation
复制标题

DOI:
10.1109/lra.2023.3303829
复制
发表时间:
2022-02
影响因子:
5.2
通讯作者:
Alireza Rezazadeh;Changhyun Choi
Alireza Rezazadeh;Changhyun Choi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Alireza Rezazadeh;Changhyun Choi

文献摘要

相似文献

以对象为中心的表示是前向预测的基本抽象。大多数现有的前向模型通过广泛的监督(例如,对象类和边界框),尽管这样的地面实况信息在现实中并不容易获得。为了解决这个问题,我们引入了KINet(关键点交互网络)-一个端到端的无监督框架,用于基于关键点表示来推理对象交互。使用视觉观察,我们的模型学习将对象与关键点坐标相关联,并发现系统的图形表示为一组关键点嵌入及其关系。然后,它使用对比估计来学习动作条件前向模型以预测未来的关键点状态。通过学习在关键点空间中执行物理推理,我们的模型自动推广到具有不同数量的对象,新颖背景和不可见对象几何形状的场景。实验表明,我们的模型在准确地执行前向预测和学习可规划的对象为中心的表示下游机器人推动操作任务的有效性。
Object-centric representation is an essential abstraction for forward prediction. Most existing forward models learn this representation through extensive supervision (e.g., object class and bounding box) although such ground-truth information is not readily accessible in reality. To address this, we introduce KINet (Keypoint Interaction Network)—an end-to-end unsupervised framework to reason about object interactions based on a keypoint representation. Using visual observations, our model learns to associate objects with keypoint coordinates and discovers a graph representation of the system as a set of keypoint embeddings and their relations. It then learns an action-conditioned forward model using contrastive estimation to predict future keypoint states. By learning to perform physical reasoning in the keypoint space, our model automatically generalizes to scenarios with a different number of objects, novel backgrounds, and unseen object geometries. Experiments demonstrate the effectiveness of our model in accurately performing forward prediction and learning plannable object-centric representations for downstream robotic pushing manipulation tasks.