UCTNet: Uncertainty-Aware Cross-Modal Transformer Network for Indoor RGB-D Semantic Segmentation

UCTNet: Uncertainty-Aware Cross-Modal Transformer Network for Indoor RGB-D Semantic Segmentation
复制标题

DOI:
10.1007/978-3-031-20056-4_2
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Xiaowen Ying;M. Chuah
Xiaowen Ying;M. Chuah
中科院分区:
其他
文献类型:
--
作者:
Xiaowen Ying;M. Chuah

文献摘要

相似文献

在本文中,我们解决了RGB-D语义切分问题。解决这一问题的关键在于1)如何从深度传感器数据中提取特征,2)如何有效地融合从两种模式提取的特征。对于第一个挑战,我们发现从传感器获得的深度信息并不总是可靠的(例如,具有反射或暗表面的物体通常具有不准确或无效的传感器读数),并且现有的使用ConvNet提取深度特征的方法没有明确地考虑不同像素位置的深度值的可靠性。为了应对这一挑战,我们提出了一种新的机制,即不确定性感知自我注意,在特征提取过程中显式地控制从不可靠的深度像素到确信的深度像素的信息流。对于第二个挑战,我们提出了一种基于交叉注意的有效且可扩展的融合模块,该模块可以在RGB和深度编码器之间进行自适应和非对称的信息交换。我们提出的框架,即UCTNet,是一个编解码器网络,它自然地结合了这两个关键设计,以实现稳健和准确的RGB-D分割。实验结果表明,UCTNet在两个RGB-D语义切分基准上的性能优于已有工作,达到了最好的性能。
In this paper, we tackle the problem of RGB-D Semantic Segmentation. The key challenges in solving this problem lie in 1) how to extract features from depth sensor data and 2) how to effectively fuse the features extracted from the two modalities. For the first challenge, we found that the depth information obtained from the sensor is not always reliable (e.g.objects with reflective or dark surfaces typically have inaccurate or void sensor readings), and existing methods that extract depth features using ConvNets did not explicitly consider the reliability of depth value at different pixel locations. To tackle this challenge, we propose a novel mechanism, namely Uncertainty-Aware Self-Attention that explicitly controls the information flow from unreliable depth pixels to confident depth pixels during feature extraction. For the second challenge, we propose an effective and scalable fusion module based on Cross-Attention that can perform adaptive and asymmetric information exchange between the RGB and depth encoder. Our proposed framework, namely UCTNet, is an encoder-decoder network that naturally incorporates these two key designs for robust and accurate RGB-D Segmentation. Experimental results show that UCTNet outperforms existing works and achieves state-of-the-art performances on two RGB-D Semantic Segmentation benchmarks.