Dynamical Scene Representation and Control with Keypoint-Conditioned Neural Radiance Field

Dynamical Scene Representation and Control with Keypoint-Conditioned Neural Radiance Field
复制标题

DOI:
10.1109/case49997.2022.9926555
复制
发表时间:
2022-08
期刊:
2022 IEEE 18th International Conference on Automation Science and Engineering (CASE)
影响因子:
--
通讯作者:
Weiyao Wang;A. S. Morgan;A. Dollar;Gregory Hager
Weiyao Wang;A. S. Morgan;A. Dollar;Gregory Hager
中科院分区:
其他
文献类型:
--
作者:
Weiyao Wang;A. S. Morgan;A. Dollar;Gregory Hager

文献摘要

相似文献

在这项工作中,我们提出了一种方法,可以学习建模动态和任意的3D场景,纯粹从2D视觉观察。我们的方法使用关键点条件神经辐射场(KP-NeRF)来捕获和建模这些场景,其总体目标是支持基于图像的机器人操作。与以前的方法不同,以前的方法通常将模型置于通用嵌入向量上进行表示,我们的隐式神经辐射函数以一组关键点为条件,这些关键点是从给定图像观察的学习编码器中推断出来的。这隐式地将可视建模组件分离为对象外观和对象姿势配置。架构中内置的这种感应偏差鼓励发现的关键点捕获机器人环境中跨时间和空间的状态转换。然后,我们学习编码关键点的前向预测模型,在关键点表示空间上构建,并执行MPC控制,以完成具有挑战性的操作任务,包括推块和关门。我们通过各种任务来评估我们的方法的性能:新颖的场景视图合成,动作条件前向预测和机器人操作任务。
In this work, we present a method that can learn to model dynamic and arbitrary 3D scenes, purely from 2D visual observations. Our approach uses a keypoint-conditioned Neural Radiance Field (KP-NeRF) to capture and model these scenes with the overarching goal of supporting image-based robot manipulation. Differentiating this from previous methods, which typically condition the model on generic embedding vectors for representation, our implicit neural radiance function is conditioned on a set of keypoints that are inferred from a learned encoder given imagery observations. This implicitly separates the visual modeling components into object appearances and object pose configurations. Such inductive bias built into the architecture encourages discovered keypoints to capture state transitions in the robot’s environment across time and space. We then learn a forward prediction model of the encoded keypoints, constructed over the keypoint representation space, and perform MPC control for challenging manipulation tasks including block pushing and door closing. We evaluate the performance of our method through various tasks: novel scene view synthesis, action-conditioned forward prediction, and robot manipulation tasks.