RGB Road Scene Material Segmentation

RGB Road Scene Material Segmentation
复制标题

DOI:
10.1007/978-3-031-26284-5_16
复制
发表时间:
2024-03
期刊:
Image Vis. Comput.
影响因子:
--
通讯作者:
Sudong Cai;Ryosuke Wakaki;S. Nobuhara;K. Nishino
Sudong Cai;Ryosuke Wakaki;S. Nobuhara;K. Nishino
中科院分区:
其他
文献类型:
--
作者:
Sudong Cai;Ryosuke Wakaki;S. Nobuhara;K. Nishino

文献摘要

相似文献

我们通过构建一个新的定制基准数据集和模型来解决RGB道路场景材质分割问题,即在纯RGB图像的真实驾驶视图中对材质进行逐像素分割。我们的新数据集KITTI-Materials基于成熟的KITTI数据集,由1000帧组成,覆盖24个不同的城市/郊区景观道路场景,为每个像素标注20种材质类别中的一种。它是第一个为现实驾驶场景中的RGB材质分割量身定制的数据集,允许我们训练和测试任何RGB材质分割模型。通过对KITTI材质模型的分析,指出纹理和上下文的提取与融合是道路场景材质外观鲁棒性的关键。我们引入道路场景素材分割网络(RMSNet),这是一个基于变形金刚的新框架,它将作为这项具有挑战性的任务的基线。RMSNet编码多尺度层次特征与自我关注。我们构造的解码器RMSNet的基础上,一种新的轻量级的自我注意力模型,我们称之为SAMixer。SAMixer实现了跨多个特征级别的信息纹理和上下文线索的自适应融合。它还通过平衡的查询关键字相似性度量显著加速了特征融合的自我注意力。我们还引入了一个内置的本地统计瓶颈,以实现进一步的效率和准确性。KITTI材料上的大量实验验证了我们的RMSNet的有效性。我们相信我们的工作为进一步研究RGB道路场景材质分割奠定了坚实的基础。
We address RGB road scene material segmentation, ie, per-pixel segmentation of materials in real-world driving views with pure RGB images, by building a new tailored benchmark dataset and model for it. Our new dataset, KITTI-Materials, based on the well-established KITTI dataset, consists of 1000 frames covering 24 different road scenes of urban/suburban landscapes, annotated with one of 20 material categories for every pixel in high quality. It is the first dataset tailored to RGB material segmentation in realistic driving scenes which allows us to train and test any RGB material segmentation model. Based on an analysis on KITTI-Materials, we identify the extraction and fusion of texture and context as the key to robust road scene material appearance. We introduce Road scene Material Segmentation Network (RMSNet), a new Transformer-based framework which will serve as a baseline for this challenging task. RMSNet encodes multi-scale hierarchical features with self-attention. We construct the decoder of RMSNet based on a novel lightweight self-attention model, which we refer to as SAMixer. SAMixer achieves adaptive fusion of informative texture and context cues across multiple feature levels. It also significantly accelerates self-attention for feature fusion with a balanced query-key similarity measure. We also introduce a built-in bottleneck of local statistics to achieve further efficiency and accuracy. Extensive experiments on KITTI-Materials validate the effectiveness of our RMSNet. We believe our work lays a solid foundation for further studies on RGB road scene material segmentation.