Robust Multimodal Depth Estimation using Transformer based Generative Adversarial Networks

Robust Multimodal Depth Estimation using Transformer based Generative Adversarial Networks
复制标题

DOI:
10.1145/3503161.3548418
复制
发表时间:
2022-10
期刊:
Proceedings of the 30th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Md Fahim Faysal Khan;Anusha Devulapally;Siddharth Advani;N. Vijaykrishnan
Md Fahim Faysal Khan;Anusha Devulapally;Siddharth Advani;N. Vijaykrishnan
中科院分区:
其他
文献类型:
--
作者:
Md Fahim Faysal Khan;Anusha Devulapally;Siddharth Advani;N. Vijaykrishnan

文献摘要

相似文献

精确测量成像传感器捕获的每个像素的绝对深度在自主导航、增强现实和机器人等实时应用中至关重要。为了预测密集深度,一般的方法是融合来自不同模态的传感器输入,例如LiDAR,相机和其他飞行时间传感器。LiDAR和其他飞行时间传感器提供准确的深度数据,但在空间和时间上都非常稀疏。为了增加缺失的深度信息,通常由于其高分辨率信息而利用RGB引导。由于依赖于多个传感器模态,鲁棒性和适应性的设计是必不可少的。在这项工作中,我们提出了一个基于transformer的自注意生成对抗网络,使用RGB和稀疏深度数据来估计密集深度。我们引入了一种新的训练方法,使模型具有鲁棒性,即使在其中一种输入方式不可用时也能工作。多头自关注机制可以动态地关注RGB图像的最显著部分或相应的稀疏深度数据,从而产生最有竞争力的结果。与其他现有的基于严重剩余连接的卷积神经网络相比,我们提出的网络还需要更少的内存用于训练和推理,使其更适合资源受限的边缘应用。源代码可在https://github.com/kocchop/robust-multimodal-fusion-gan上获得
Accurately measuring the absolute depth of every pixel captured by an imaging sensor is of critical importance in real-time applications such as autonomous navigation, augmented reality and robotics. In order to predict dense depth, a general approach is to fuse sensor inputs from different modalities such as LiDAR, camera and other time-of-flight sensors. LiDAR and other time-of-flight sensors provide accurate depth data but are quite sparse, both spatially and temporally. To augment missing depth information, generally RGB guidance is leveraged due to its high resolution information. Due to the reliance on multiple sensor modalities, design for robustness and adaptation is essential. In this work, we propose a transformer-like self-attention based generative adversarial network to estimate dense depth using RGB and sparse depth data. We introduce a novel training recipe for making the model robust so that it works even when one of the input modalities is not available. The multi-head self-attention mechanism can dynamically attend to most salient parts of the RGB image or corresponding sparse depth data producing the most competitive results. Our proposed network also requires less memory for training and inference compared to other existing heavily residual connection based convolutional neural networks, making it more suitable for resource-constrained edge applications. The source code is available at: https://github.com/kocchop/robust-multimodal-fusion-gan