An Iterative Semi-supervised Approach with Pixel-wise Contrastive Loss for Road Extraction in Aerial Images

An Iterative Semi-supervised Approach with Pixel-wise Contrastive Loss for Road Extraction in Aerial Images
复制标题

DOI:
10.1145/3606374
复制
发表时间:
2023-07
期刊:
ACM Transactions on Multimedia Computing, Communications and Applications
影响因子:
--
通讯作者:
Huijie Zhang;Pu Li;Xiaobai Liu;Xianfeng Yang;Li An
Huijie Zhang;Pu Li;Xiaobai Liu;Xianfeng Yang;Li An
中科院分区:
其他
文献类型:
--
作者:
Huijie Zhang;Pu Li;Xiaobai Liu;Xianfeng Yang;Li An

文献摘要

相似文献

航空图像中的道路提取在人工智能和多媒体计算中有着广泛的应用,包括交通模式分析和停车位规划。学习深度神经网络虽然非常成功,但需要大量高质量的注释,而获取这些注释既耗时又昂贵。在这项工作中,我们提出了一种基于图像的半监督道路提取方法,其中只有一小部分标记图像可用于训练,以应对这一挑战。我们设计了一个像素级的对比损失来对网络训练进行自我监督,以利用大量的未标记图像。其关键思想是识别重叠图像区域对(正)或非重叠图像区域对(负),并鼓励网络对正图像对做出相似的输出或对负图像对做出不相似的输出。我们还开发了一种负采样策略来过滤过程中的假阴性样本。介绍了一种迭代过程,将网络应用于原始图像,以生成伪标签,过滤并选择具有所提出的对比损失的高质量标签,并用扩大的训练数据集重新训练网络。我们重复这些迭代步骤,直到收敛。我们在公共的SpaceNet3和DeepGlobe Road数据集上进行了大量的实验,验证了所提方法的有效性。结果表明,我们提出的方法在公共图像分割基准上达到了最先进的结果,并且显著优于其他半监督方法。
Extracting roads in aerial images has numerous applications in artificial intelligence and multimedia computing, including traffic pattern analysis and parking space planning. Learning deep neural networks, though very successful, demand vast amounts of high-quality annotations, of which acquisition is time-consuming and expensive. In this work, we propose a semi-supervised approach for image-based road extraction in which only a small set of labeled images are available for training to address this challenge. We design a pixel-wise contrastive loss to self-supervise the network training to utilize the large corpus of unlabeled images. The key idea is to identify pairs of overlapping image regions (positive) or non-overlapping image regions (negative) and encourage the network to make similar outputs for positive pairs or dissimilar outputs for negative pairs. We also develop a negative sampling strategy to filter false-negative samples during the process. An iterative procedure is introduced to apply the network over raw images to generate pseudo-labels, filter and select high-quality labels with the proposed contrastive loss, and retrain the network with the enlarged training dataset. We repeat these iterative steps until convergence. We validate the effectiveness of the proposed methods by performing extensive experiments on the public SpaceNet3 and DeepGlobe Road datasets. Results show that our proposed method achieves state-of-the-art results on public image segmentation benchmarks and significantly outperforms other semi-supervised methods.