Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation

Act3D: 3D Feature Field Transformers for Multi-Task Robotic Manipulation
复制标题

DOI:
--
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Théophile Gervet;Zhou Xian;N. Gkanatsios;Katerina Fragkiadaki
Théophile Gervet;Zhou Xian;N. Gkanatsios;Katerina Fragkiadaki
中科院分区:
其他
文献类型:
--
作者:
Théophile Gervet;Zhou Xian;N. Gkanatsios;Katerina Fragkiadaki

文献摘要

相似文献

3D感知表示非常适合机器人操作,因为它们很容易对遮挡进行编码,并简化空间推理。许多操纵任务对末端执行器姿态预测的空间精度要求很高,这通常需要高分辨率的3D特征网格,而这些网格的计算代价很高。因此,大多数操作策略直接在2D中操作,而不是3D诱导偏差。在本文中,我们介绍了Act3D,一个操作策略转换器,它使用一个3D特征域来表示机器人的工作空间,该特征域具有根据手头任务的自适应分辨率。该模型利用感知深度将预先训练好的2D特征提升到3D,并关注它们来计算采样的3D点的特征。它以从粗到精的方式对三维点网格进行采样,使用相对位置注意力对其进行特征化,并选择下一轮点采样的焦点位置。通过这种方式,它可以高效地计算出高空间分辨率的3D动作图。Act3D在RL-BENCH中设置了一个新的最先进的操作基准,在74个RLBch任务上,它比以前的SOTA 2D多视图策略实现了10%的绝对改进,与以前的SOTA 3D策略相比,计算量减少了3倍,实现了22%的绝对改进。在消融实验中,我们量化了相对空间注意力、大规模视觉语言预先训练的2D主干和从粗到精的注意的权重挂钩的重要性。代码和视频可在我们的项目网站上找到:https://act3d.github.io/.
3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning. Many manipulation tasks require high spatial precision in end-effector pose prediction, which typically demands high-resolution 3D feature grids that are computationally expensive to process. As a result, most manipulation policies operate directly in 2D, foregoing 3D inductive biases. In this paper, we introduce Act3D, a manipulation policy transformer that represents the robot's workspace using a 3D feature field with adaptive resolutions dependent on the task at hand. The model lifts 2D pre-trained features to 3D using sensed depth, and attends to them to compute features for sampled 3D points. It samples 3D point grids in a coarse to fine manner, featurizes them using relative-position attention, and selects where to focus the next round of point sampling. In this way, it efficiently computes 3D action maps of high spatial resolution. Act3D sets a new state-of-the-art in RL-Bench, an established manipulation benchmark, where it achieves 10% absolute improvement over the previous SOTA 2D multi-view policy on 74 RLBench tasks and 22% absolute improvement with 3x less compute over the previous SOTA 3D policy. We quantify the importance of relative spatial attention, large-scale vision-language pre-trained 2D backbones, and weight tying across coarse-to-fine attentions in ablative experiments. Code and videos are available on our project website: https://act3d.github.io/.