Spatio-Temporal Fusion of LiDAR and Camera Data for Omnidirectional Depth Perception

Spatio-Temporal Fusion of LiDAR and Camera Data for Omnidirectional Depth Perception
复制标题

DOI:
10.1177/03611981231184187
复制
发表时间:
2023-07
影响因子:
1.7
通讯作者:
Linlin Zhang;Xiang Yu;Y. Adu-Gyamfi;Carlos C. Sun
Linlin Zhang;Xiang Yu;Y. Adu-Gyamfi;Carlos C. Sun
中科院分区:
工程技术4区
文献类型:
--
作者:
Linlin Zhang;Xiang Yu;Y. Adu-Gyamfi;Carlos C. Sun

文献摘要

相似文献

物体识别和深度感知是两个紧密耦合的任务,是必不可少的态势感知。大多数自主系统能够通过处理和整合来自各种传感器的数据流来执行这些任务。操作这些系统所需的多个硬件和复杂的软件架构使得它们的扩展和操作成本高昂。本文实现了一个快速的,单目视觉系统,可用于同时对象识别和深度感知。我们借用了最先进的对象识别系统YOLOv3的架构,并通过合并距离和修改其损失函数和预测向量来扩展其架构,使其能够在两个任务上进行多任务处理。视觉系统在通过LiDAR测量与互补360度相机耦合获得的大型数据库上进行训练,以生成高保真标记数据集。多用途网络的性能是在一个测试数据集上进行评估的,该测试数据集由在不同道路网络上收集的总共7,634个对象组成。与地面实况激光雷达数据相比,所提出的网络在10 m以内的乘用车上实现了11%的平均绝对百分比误差率,在10 m以内和10 m以外的卡车上分别实现了7%或9%的平均误差率。研究还发现,在建模网络中添加第二个任务(深度感知)可以将目标检测的准确性提高约3%。所提出的多用途模型可用于开发自动报警系统、交通监控和安全监控。
Object recognition and depth perception are two tightly coupled tasks that are indispensable for situational awareness. Most autonomous systems are able to perform these tasks by processing and integrating data streaming from a variety of sensors. The multiple hardware and sophisticated software architectures required to operate these systems makes them expensive to scale and operate. This paper implements a fast, monocular vision system that can be used for simultaneous object recognition and depth perception. We borrow from the architecture of a start-of-the-art object recognition system, YOLOv3, and extend its architecture by incorporating distances and modifying its loss functions and prediction vectors to enable it to multitask on both tasks. The vision system is trained on a large database acquired through the coupling of LiDAR measurements with complementary 360-degree camera to generate a high-fidelity labeled dataset. The performance of the multipurpose network is evaluated on a test dataset consisting of a total of 7,634 objects collected on a different road network. When compared with ground truth LiDAR data, the proposed network achieves a mean absolute percentage error rate of 11% on the passenger car within 10 m and a mean error rate of 7% or 9% on the truck within 10 m and beyond 10 m, respectively. It was also observed that adding a second task (depth perception) to the modeling network improved the accuracy of object detection by about 3%. The proposed multipurpose model can be used for the development of automated alert systems, traffic monitoring, and safety monitoring.