Fast Monocular Depth Estimation via Side Prediction Aggregation with Continuous Spatial Refinement

Fast Monocular Depth Estimation via Side Prediction Aggregation with Continuous Spatial Refinement
复制标题

DOI:
10.1109/tmm.2021.3140001
复制
发表时间:
2023
影响因子:
7.3
通讯作者:
Jipeng Wu;R. Ji;Qiang Wang;Shengchuan Zhang;Xiaoshuai Sun;Yan Wang;Mingliang Xu;Feiyue Huang
Jipeng Wu;R. Ji;Qiang Wang;Shengchuan Zhang;Xiaoshuai Sun;Yan Wang;Mingliang Xu;Feiyue Huang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jipeng Wu;R. Ji;Qiang Wang;Shengchuan Zhang;Xiaoshuai Sun;Yan Wang;Mingliang Xu;Feiyue Huang

文献摘要

相似文献

最近的工作已经验证了将空间信息集成到深度网络中来改进像素级预测任务的好处,例如单目深度估计。然而,如何有效和稳健地整合空间线索仍然是一个悬而未决的问题。在本文中,我们引入了边预测聚集(简称SPA)方法来增强场景结构信息从低层到高层的嵌入。为了提高估计精度,该方法在不增加额外计算量的情况下,进一步增加了多分辨率下的连续空间精化损失(SRL)。此外,所提出的序列网络还可以在多分辨率下进行对抗性学习。这种对抗性求精策略大大提高了深度估计的精度,只需少量的额外计算。在不使用任何预训练模型的情况下,我们的网络在Kitti、NYUD V2和CitySees数据集上达到了最先进的精度,从而实现了在线实时深度估计。
Recent works have validated the benefit of integrating spatial information into deep networks to improve pixel-level prediction tasks such as monocular depth estimation. However, how to efficiently and robustly integrate spatial cues retains as an open problem. In this paper, we introduce the Side Prediction Aggregation (termed SPA) method to enhance the embedding of scene structural information from low-level to high-level layers. To improve the estimation accuracy, the proposed method is further equipped with continuous Spatial Refinement Loss (termed SRL) at multiple resolutions with negligible extra computation. Besides, the proposed sequential network can further perform adversarial learning at multiple resolutions. Such an adversarial refinement strategy greatly improves the accuracy of estimated depth with a little extra computation. Without using any pre-trained models, our network achieves the the-state-of-art accuracy on KITTI, NYUD V2, and Cityscapes datasets, which has achieved real-time depth estimation online.