Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement

Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
复制标题

DOI:
10.15607/rss.2023.xix.030
复制
发表时间:
2023-04
期刊:
Robotics: Science and Systems XIX
影响因子:
--
通讯作者:
N. Gkanatsios;Ayush Jain;Zhou Xian;Yunchu Zhang;C. Atkeson;Katerina Fragkiadaki
N. Gkanatsios;Ayush Jain;Zhou Xian;Yunchu Zhang;C. Atkeson;Katerina Fragkiadaki
中科院分区:
其他
文献类型:
--
作者:
N. Gkanatsios;Ayush Jain;Zhou Xian;Yunchu Zhang;C. Atkeson;Katerina Fragkiadaki

文献摘要

被引文献

相似文献

语言是合成的;指令可以表达要在场景中的对象之间保持的多个关系约束,机器人被分派任务来重新布置该对象。我们在这项工作中的重点是一个可扩展的场景重排框架,概括为更长的指令和空间概念的组成,从来没有见过在培训时间。我们建议表示语言指导的空间概念的能量函数相对对象的安排。语言解析器将指令映射到相应的能量函数,开放词汇的视觉语言模型将它们的参数与场景中的相关对象联系起来。我们生成目标场景配置的能量函数,每个语言谓词的指令的总和梯度下降。然后,本地基于视觉的策略将对象重新定位到推断的目标位置。我们测试我们的模型建立的预防指导的操作基准,以及基准的组合指令,我们介绍。我们表明,我们的模型可以执行高度组合指令零射击在模拟和真实的世界。它比语言到动作反应策略和大型语言模型规划器的性能要好得多,特别是对于涉及多个空间概念组合的长指令。模拟和真实世界的机器人执行视频,以及我们的代码和数据集可在我们的网站上公开获取:https://ebmplanner.github.io。
Language is compositional; an instruction can express multiple relation constraints to hold among objects in a scene that a robot is tasked to rearrange. Our focus in this work is an instructable scene-rearranging framework that generalizes to longer instructions and to spatial concept compositions never seen at training time. We propose to represent language-instructed spatial concepts with energy functions over relative object arrangements. A language parser maps instructions to corresponding energy functions and an open-vocabulary visual-language model grounds their arguments to relevant objects in the scene. We generate goal scene configurations by gradient descent on the sum of energy functions, one per language predicate in the instruction. Local vision-based policies then re-locate objects to the inferred goal locations. We test our model on established instruction-guided manipulation benchmarks, as well as benchmarks of compositional instructions we introduce. We show our model can execute highly compositional instructions zero-shot in simulation and in the real world. It outperforms language-to-action reactive policies and Large Language Model planners by a large margin, especially for long instructions that involve compositions of multiple spatial concepts. Simulation and real-world robot execution videos, as well as our code and datasets are publicly available on our website: https://ebmplanner.github.io.