Focal-WNet: An Architecture Unifying Convolution and Attention for Depth Estimation

Focal-WNet: An Architecture Unifying Convolution and Attention for Depth Estimation
复制标题

Focal-WNet:一种统一卷积和注意力来进行深度估计的架构

DOI:
10.1109/i2ct54291.2022.9824488
复制
发表时间:
2022
期刊:
2022 IEEE 7th International conference for Convergence in Technology (I2CT)
影响因子:
--
通讯作者:
J. Swaminathan
J. Swaminathan
中科院分区:
--
文献类型:
--
作者:
G. Manimaran;J. Swaminathan

文献摘要

被引文献

相似文献

从单幅RGB图像中提取深度信息是计算机视觉中一项具有广泛应用前景的基础性和挑战性任务。这项任务不能用多视几何等传统方法来解决,只能通过深度学习来解决。现有的卷积神经网络方法由于缺乏长程相关性,导致结果不一致和模糊。随着变形器网络在计算机视觉领域取得的成功,我们利用这一思想提出了一种新的体系结构,称为Focus-WNET。该体系结构由两个独立的编码器和一个解码器组成。该网络的主要目标是学习大多数单目深度线索,如相对尺度、对比度差异、纹理梯度等。我们结合焦点自我注意而不是普通的自我注意来降低网络的计算复杂度。除了焦点转换器层,我们利用卷积结构来学习深度提示,而这些提示不能由变压器单独学习,因为一些提示,如遮挡,需要局部接受场,更容易被转换网络学习。大量的实验表明,提出的FOCUS-WNET在两个具有挑战性的数据集上取得了竞争的结果。我们的代码和预先训练的重量可在https://github.com/Goubeast/Focal-WNet上获得
Extracting depth information from a single RGB image is a fundamental and challenging task in computer vision with wide-ranging applications. This task cannot be solved using traditional methods like multi-view geometry but only by deep learning. Existing methods using convolutional neural nets produce inconsistent and blurry results due to the lack of long-range dependencies. With the recent success of Transformer networks in computer vision, which can process information locally and globally, we leverage this idea to propose a novel architecture named Focal-WNet in this paper. This architecture consists of two separate encoders and a single decoder. The main aim of this network is to learn most monocular depth cues like relative scale, contrast differences, texture gradient, etc. We incorporate focal self-attention instead of vanilla self-attention to reduce the computational complexity of the network. Along with the focal transformer layers, we leverage a convolutional architecture to learn depth cues that cannot be learned by a transformer alone as some cues like occlusion require a local receptive field and are easier for a conv-net to learn. Extensive experiments show that the proposed Focal-WNet achieves competitive results on two challenging datasets. Our code and pre-trained weights are available at https://github.com/Goubeast/Focal-WNet