VIOLA: Imitation Learning for Vision-Based Manipulation with Object Proposal Priors

VIOLA: Imitation Learning for Vision-Based Manipulation with Object Proposal Priors
复制标题

DOI:
10.48550/arxiv.2210.11339
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yifeng Zhu;Abhishek Joshi;P. Stone;Yuke Zhu
Yifeng Zhu;Abhishek Joshi;P. Stone;Yuke Zhu
中科院分区:
其他
文献类型:
--
作者:
Yifeng Zhu;Abhishek Joshi;P. Stone;Yuke Zhu

文献摘要

相似文献

我们介绍VIOLA,一个以对象为中心的模仿学习方法来学习机器人操作的闭环视觉策略。我们的方法基于预训练的视觉模型的一般对象建议构建以对象为中心的表示。VIOLA使用基于transformer的策略来推理这些表示,并关注与任务相关的视觉因素以进行动作预测。这种基于对象的结构先验提高了深度模仿学习算法对对象变化和环境扰动的鲁棒性。我们定量评估VIOLA在模拟和真实的机器人。VIOLA的成功率比最先进的模仿学习方法高出45.8\%$。它还成功地部署在物理机器人上,以解决具有挑战性的长期任务,例如餐桌布置和咖啡制作。更多视频和模型细节可以在补充材料和项目网站https://ut-austin-rpl.github.io/VIOLA中找到。
We introduce VIOLA, an object-centric imitation learning approach to learning closed-loop visuomotor policies for robot manipulation. Our approach constructs object-centric representations based on general object proposals from a pre-trained vision model. VIOLA uses a transformer-based policy to reason over these representations and attend to the task-relevant visual factors for action prediction. Such object-based structural priors improve deep imitation learning algorithm's robustness against object variations and environmental perturbations. We quantitatively evaluate VIOLA in simulation and on real robots. VIOLA outperforms the state-of-the-art imitation learning methods by $45.8\%$ in success rate. It has also been deployed successfully on a physical robot to solve challenging long-horizon tasks, such as dining table arrangement and coffee making. More videos and model details can be found in supplementary material and the project website: https://ut-austin-rpl.github.io/VIOLA .