Dynamic-Resolution Model Learning for Object Pile Manipulation

Dynamic-Resolution Model Learning for Object Pile Manipulation
复制标题

DOI:
10.48550/arxiv.2306.16700
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yixuan Wang;Yunzhu Li;K. Driggs-Campbell;Li Fei-Fei-Li-Fei-Fei-48004138;Jiajun Wu
Yixuan Wang;Yunzhu Li;K. Driggs-Campbell;Li Fei-Fei-Li-Fei-Fei-48004138;Jiajun Wu
中科院分区:
其他
文献类型:
--
作者:
Yixuan Wang;Yunzhu Li;K. Driggs-Campbell;Li Fei-Fei-Li-Fei-Fei-48004138;Jiajun Wu

文献摘要

被引文献

相似文献

从视觉观察中学习到的动力学模型已在各种机器人操作任务中显示出有效性。学习此类动力学模型的关键问题之一是使用何种场景表示。先前的工作通常假定采用固定维度或分辨率的表示,这对于简单任务可能效率低下,对于更复杂的任务可能无效。在这项工作中,我们研究如何在不同的抽象层次上学习动态和自适应的表示,以在效率和有效性之间实现最佳平衡。具体而言,我们构建环境的动态分辨率粒子表示,并使用图神经网络(GNNs)学习一个统一的动力学模型,该模型允许连续选择抽象层次。在测试阶段,智能体可以在每个模型预测控制(MPC)步骤中自适应地确定最佳分辨率。我们在物体堆操作中评估我们的方法,这是我们在烹饪、农业、制造业和制药应用中常见的任务。通过在模拟和现实世界中的综合评估,我们表明,在收集、分类和重新分配由咖啡豆、杏仁、玉米等各种实例组成的颗粒状物体堆时,我们的方法比最先进的固定分辨率基线取得了显著更好的性能。
Dynamics models learned from visual observations have shown to be effective in various robotic manipulation tasks. One of the key questions for learning such dynamics models is what scene representation to use. Prior works typically assume representation at a fixed dimension or resolution, which may be inefficient for simple tasks and ineffective for more complicated tasks. In this work, we investigate how to learn dynamic and adaptive representations at different levels of abstraction to achieve the optimal trade-off between efficiency and effectiveness. Specifically, we construct dynamic-resolution particle representations of the environment and learn a unified dynamics model using graph neural networks (GNNs) that allows continuous selection of the abstraction level. During test time, the agent can adaptively determine the optimal resolution at each model-predictive control (MPC) step. We evaluate our method in object pile manipulation, a task we commonly encounter in cooking, agriculture, manufacturing, and pharmaceutical applications. Through comprehensive evaluations both in the simulation and the real world, we show that our method achieves significantly better performance than state-of-the-art fixed-resolution baselines at the gathering, sorting, and redistribution of granular object piles made with various instances like coffee beans, almonds, corn, etc.