Learning Scene Dynamics from Point Cloud Sequences

Learning Scene Dynamics from Point Cloud Sequences
复制标题

DOI:
10.1007/s11263-021-01551-y
复制
发表时间:
2022-01-23
影响因子:
19.5
通讯作者:
Rangarajan, Anand
Rangarajan, Anand
中科院分区:
计算机科学2区
文献类型:
--
作者:
He, Pan;Emami, Patrick;Rangarajan, Anand

文献摘要

被引文献

相似文献

理解3D场景是自主代理的关键先决条件。最近,LiDAR和其他传感器以点云帧的时间序列的形式提供了大量数据。在这项工作中,我们提出了一个新的问题顺序场景流估计(SSFE),其目的是预测三维场景流的所有对点云在一个给定的序列。这与先前研究的场景流估计问题不同,场景流估计问题集中在两个帧上。我们引入了SPCM-Net架构,该架构通过计算相邻点云之间的多尺度时空相关性,然后用顺序不变的递归单元聚合时间上的相关性来解决这个问题。我们的实验评估证实,与仅使用两个帧相比,点云序列的递归处理导致显著更好的SSFE。此外,我们证明,这种方法可以有效地修改为连续点云预测(SPF),一个相关的问题,需要预测未来的点云帧。我们的实验结果进行评估,使用一个新的基准SSFE和SPF合成和真实的数据集组成。以前,用于场景流估计的数据集被限制为两帧。我们为这些数据集提供了非平凡的扩展,用于多帧估计和预测。由于难以获得真实世界数据集的地面真实运动,我们使用自监督训练和评估指标。我们相信,这一基准将是在这一领域的未来研究的关键。所有基准测试和模型的代码都可以在(https://github.com/BestSonny/SPCM)上访问。
Understanding 3D scenes is a critical prerequisite for autonomous agents. Recently, LiDAR and other sensors have made large amounts of data available in the form of temporal sequences of point cloud frames. In this work, we propose a novel problem-sequential scene flow estimation (SSFE)-that aims to predict 3D scene flow for all pairs of point clouds in a given sequence. This is unlike the previously studied problem of scene flow estimation which focuses on two frames. We introduce the SPCM-Net architecture, which solves this problem by computing multi-scale spatiotemporal correlations between neighboring point clouds and then aggregating the correlation across time with an order-invariant recurrent unit. Our experimental evaluation confirms that recurrent processing of point cloud sequences results in significantly better SSFE compared to using only two frames. Additionally, we demonstrate that this approach can be effectively modified for sequential point cloud forecasting (SPF), a related problem that demands forecasting future point cloud frames. Our experimental results are evaluated using a new benchmark for both SSFE and SPF consisting of synthetic and real datasets. Previously, datasets for scene flow estimation have been limited to two frames. We provide non-trivial extensions to these datasets for multi-frame estimation and prediction. Due to the difficulty of obtaining ground truth motion for real-world datasets, we use self-supervised training and evaluation metrics. We believe that this benchmark will be pivotal to future research in this area. All code for benchmark and models will be made accessible at (https://github.com/BestSonny/SPCM).