An Efficient Accelerator for Deep Learning-based Point Cloud Registration on FPGAs

An Efficient Accelerator for Deep Learning-based Point Cloud Registration on FPGAs
复制标题

DOI:
10.1109/pdp59025.2023.00018
复制
发表时间:
2022-03
期刊:
2023 31st Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP)
影响因子:
--
通讯作者:
K. Sugiura;Hiroki Matsutani
K. Sugiura;Hiroki Matsutani
中科院分区:
其他
文献类型:
--
作者:
K. Sugiura;Hiroki Matsutani

文献摘要

相似文献

点云配准是许多机器人应用的基础,例如里程计和同时定位和地图绘制(SLAM),其对于自主移动的机器人越来越重要。这种机器人上的计算资源和功率预算的限制促使我们研究低成本边缘设备上的资源有效的注册方法。在本文中,我们提出了一种基于FPGA的新型3D点云配准管道,该管道基于最近的基于深度学习的方法PointNetLK。基于分析结果,我们专注于PointNet特征提取,因为它成为一个主要的瓶颈;我们通过以流水线方式逐个消耗每个输入点,而不是一次处理整个点云,来提高其可扩展性和内存效率。然后,我们设计了一个完全并行化和流水线加速器,包括一个定制的PointNet IP核,它适合低成本和中档FPGA(例如,Avnet Ultra 96 v2和Xilinx ZCU 104)。实验结果表明,我们提出的流水线实现了高达21.34倍和69.60倍的注册速度比香草PointNetLK和ICP,分别,而仅消耗722 mW,并保持相同的精度水平。
Point cloud registration is the basis for many robotic applications such as odometry and Simultaneous Localization And Mapping (SLAM), which are increasingly important for autonomous mobile robots. The limitation of computational resources and power budgets on such robots motivates us to study the resource-efficient registration method on low-cost edge devices. In this paper, we propose an FPGA-based novel pipeline for 3D point cloud registration built upon a recent deep learning-based method, PointNetLK. Based on the profiling results, we focus on the PointNet feature extraction as it becomes a major bottleneck; we improve its scalability and memory-efficiency by consuming each input point one-by-one in a pipelined manner instead of processing the whole point cloud at once. We then design a fully-parallelized and pipelined accelerator consisting of a custom PointNet IP core, which fits within both low-cost and mid-range FPGAs (e.g., Avnet Ultra96v2 and Xilinx ZCU104). Experimental results show that our proposed pipeline achieves up to 21.34x and 69.60x faster registration speed than the vanilla PointNetLK and ICP, respectively, while only consuming 722mW and maintaining the same level of accuracy.