Uncertainty from Motion for DNN Monocular Depth Estimation

Uncertainty from Motion for DNN Monocular Depth Estimation
复制标题

DOI:
10.1109/icra46639.2022.9812222
复制
发表时间:
2022-05
期刊:
2022 International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Soumya Sudhakar;V. Sze;S. Karaman
Soumya Sudhakar;V. Sze;S. Karaman
中科院分区:
其他
文献类型:
--
作者:
Soumya Sudhakar;V. Sze;S. Karaman

文献摘要

相似文献

在资源受限的平台上部署用于安全关键场景中的单目深度估计的深度神经网络(DNN)需要良好校准和有效的不确定性估计。然而,许多流行的不确定性估计技术,包括最新的集成和流行的基于采样的方法,需要每个输入都有多个推断,这使得它们很难在延迟受限或能量受限的场景中部署。我们提出了一种新的算法,称为运动不确定性(UfM),它只需要每个输入一次推理。UFM通过增量地合并视频序列中在多个视图中看到的点的每像素深度预测和每像素任意不确定性预测来利用视频输入中的时间冗余。当UfM应用于系综时,我们证明了UfM可以通过在每一帧上只运行一个系综成员并融合帧序列上的不确定性来在一小部分能量下保持系综的不确定性质量。在一组使用FCDenseNet和八个未分发和未分发视频序列的代表性实验中,UfM提供了与大小为10的合奏相当的不确定性质量,同时仅消耗该合奏11.3%的能量,并且在单个NVIDIA RTX 2080 Ti GPU上运行速度快6.4倍,为资源受限的实时场景提供了接近合奏的不确定性质量。
Deployment of deep neural networks (DNNs) for monocular depth estimation in safety-critical scenarios on resource-constrained platforms requires well-calibrated and efficient uncertainty estimates. However, many popular uncertainty estimation techniques, including state-of-the-art ensembles and popular sampling-based methods, require multiple inferences per input, making them difficult to deploy in latency-constrained or energy-constrained scenarios. We propose a new algorithm, called Uncertainty from Motion (UfM), that requires only one inference per input. UfM exploits the temporal redundancy in video inputs by merging incrementally the per-pixel depth prediction and per-pixel aleatoric uncertainty prediction of points that are seen in multiple views in the video sequence. When UfM is applied to ensembles, we show that UfM can retain the uncertainty quality of ensembles at a fraction of the energy by running only a single ensemble member at each frame and fusing the uncertainty over the sequence of frames. In a set of representative experiments using FCDenseNet and eight indistribution and out-of-distribution video sequences, UfM offers comparable uncertainty quality to an ensemble of size 10 while consuming only 11.3% of the ensemble's energy and running 6.4× faster on a single Nvidia RTX 2080 Ti GPU, enabling near ensemble uncertainty quality for resource-constrained, real-time scenarios.