Depth Monocular Estimation with Attention-based Encoder-Decoder Network from Single Image

Depth Monocular Estimation with Attention-based Encoder-Decoder Network from Single Image
复制标题

DOI:
10.1109/hpcc-dss-smartcity-dependsys57074.2022.00271
复制
发表时间:
2022-10
期刊:
2022 IEEE 24th Int Conf on High Performance Computing & Communications; 8th Int Conf on Data Science & Systems; 20th Int Conf on Smart City; 8th Int Conf on Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC/DSS/SmartCity/DependSys)
影响因子:
--
通讯作者:
Xin Zhang;R. Abdelfattah;Yuqi Song;Samuel A. Dauchert;Xiaofen Wang
Xin Zhang;R. Abdelfattah;Yuqi Song;Samuel A. Dauchert;Xiaofen Wang
中科院分区:
其他
文献类型:
--
作者:
Xin Zhang;R. Abdelfattah;Yuqi Song;Samuel A. Dauchert;Xiaofen Wang

文献摘要

相似文献

深度信息是感知的基础,对于自动驾驶、机器人和其他资源受限的应用至关重要。准确、有效地获取深度信息,可以在动态环境中快速响应。使用LIDAR和RADAR的基于传感器的方法以高功耗、价格和体积为代价获得高精度。虽然由于深度学习的进步,基于视觉的方法最近受到了很多关注,并且可以克服这些缺点。在这项工作中,我们探索了一个极端的情况下,基于视觉的设置:估计深度图从一个单目图像严重困扰网格文物和模糊的边缘。为了解决这个问题,我们首先设计了一个卷积注意力机制块(CAMB),它依次由通道注意力和空间注意力组成,并将这些CAMB插入到跳过连接中。因此,我们的新方法可以找到当前图像的焦点,以最小的开销,并避免损失的深度功能。然后,结合深度值、X轴、Y轴和对角线方向的梯度以及结构相似性指数测度(SSIM),我们提出了新的损失函数。此外,我们利用像素块来加速损失函数的计算。最后,我们通过对两个大规模图像数据集,即KITTI和NYU-V2的综合实验表明,我们的方法优于几个代表性的基线。
Depth information is the foundation of perception, essential for autonomous driving, robotics, and other source- constrained applications. Promptly obtaining accurate and effi-cient depth information allows for a rapid response in dynamic environments. Sensor-based methods using LIDAR and RADAR obtain high precision at the cost of high power consumption, price, and volume. While due to advances in deep learning, vision-based approaches have recently received much attention and can overcome these drawbacks. In this work, we explore an extreme scenario in vision-based settings: estimate a depth map from one monocular image severely plagued by grid artifacts and blurry edges. To address this scenario, We first design a convolutional attention mechanism block (CAMB) which consists of channel attention and spatial attention sequentially and insert these CAMBs into skip connections. As a result, our novel approach can find the focus of current image with minimal overhead and avoid losses of depth features. Next, by combining the depth value, the gradients of X axis, Y axis and diagonal directions, and the structural similarity index measure (SSIM), we propose our novel loss function. Moreover, we utilize pixel blocks to accelerate the computation of the loss function. Finally, we show, through comprehensive experiments on two large- scale image datasets, i.e. KITTI and NYU - V2, that our method outperforms several representative baselines.