Self-Supervised Correspondence in Visuomotor Policy Learning

Self-Supervised Correspondence in Visuomotor Policy Learning
复制标题

DOI:
10.1109/lra.2019.2956365
复制
发表时间:
2019-09
影响因子:
5.2
通讯作者:
Peter R. Florence;Lucas Manuelli;Russ Tedrake
Peter R. Florence;Lucas Manuelli;Russ Tedrake
中科院分区:
计算机科学2区
文献类型:
--
作者:
Peter R. Florence;Lucas Manuelli;Russ Tedrake

文献摘要

被引文献

相似文献

在这封信中,我们探索了使用自监督对应来提高视觉运动策略学习的泛化性能和样本效率。先前的工作主要使用自动编码、基于姿态的损失和端到端策略优化等方法来训练视觉运动策略的视觉部分。相反,我们提出了一种使用自监督密集视觉对应训练的方法,并表明这使得视觉运动策略学习在少量数据下具有惊人的高泛化性能。使用模仿学习,我们在具有挑战性的操作任务上展示了广泛的硬件验证,只需50个演示。我们学习的策略可以泛化对象的类别,对可变形的对象配置做出反应,并在各种背景下操作无纹理的对称对象,所有这些都是闭环的,基于实时视觉的策略。模拟模仿学习实验表明,与自动编码和端到端训练相比,对应训练具有样本复杂性和泛化优势。
In this letter, we explore using self-supervised correspondence for improving the generalization performance and sample efficiency of visuomotor policy learning. Prior work has primarily used approaches such as autoencoding, pose-based losses, and end-to-end policy optimization in order to train the visual portion of visuomotor policies. We instead propose an approach using self-supervised dense visual correspondence training and show that this enables visuomotor policy learning with surprisingly high generalization performance with modest amounts of data. Using imitation learning, we demonstrate extensive hardware validation on challenging manipulation tasks with as few as 50 demonstrations. Our learned policies can generalize across classes of objects, react to deformable object configurations, and manipulate textureless symmetrical objects in a variety of backgrounds, all with closed-loop, real-time vision-based policies. Simulated imitation learning experiments suggest that correspondence training offers sample complexity and generalization benefits compared to autoencoding and end-to-end training.