Learning Aerial Image Segmentation From Online Maps

Learning Aerial Image Segmentation From Online Maps
复制标题

DOI:
10.1109/tgrs.2017.2719738
复制
发表时间:
2017-11-01
影响因子:
8.2
通讯作者:
Schindler, Konrad
Schindler, Konrad
中科院分区:
工程技术1区
文献类型:
--
作者:
Kaiser, Pascal;Wegner, Jan Dirk;Schindler, Konrad

文献摘要

被引文献

相似文献

本文研究高分辨率(航空)图像的语义分割,通过监督分类为每个像素分配一个语义类别标签,作为自动地图生成的基础。最近,深度卷积神经网络(CNN)表现出了令人印象深刻的性能,并迅速成为语义分割的事实上的标准,另外,不再需要特定任务的特征设计。然而,深度学习方法的一个主要缺点是它们非常需要数据,从而加剧了监督分类的长期瓶颈,无法获得足够的标注训练数据。另一方面,已经观察到它们对训练标签中的噪声具有相当强的鲁棒性。这开启了一种有趣的可能性,即避免注释大量的训练数据,而是从现有的遗留数据或可能显示高噪声水平的众包地图中训练分类器。本文讨论的问题是:使用大规模公开可用的标签进行培训,是否可以取代大部分的人工标签工作,并仍然取得足够的性能?这样的数据将不可避免地包含相当大一部分错误,但作为回报,世界上更大的地区几乎可以获得无限数量的数据。采用最新的CNN框架对航拍图像中的建筑物和道路进行语义分割,并比较了其在使用不同训练数据集时的性能,从人工标记的同一城市像素精确的地面真实数据到从远距离位置的OpenStreetMap数据派生的自动训练数据。我们报告的结果表明,通过利用噪声较大的大规模训练数据,可以显著减少人工标注工作量,获得令人满意的性能。
This paper deals with semantic segmentation of high-resolution (aerial) images where a semantic class label is assigned to each pixel via supervised classification as a basis for automatic map generation. Recently, deep convolutional neural networks (CNNs) have shown impressive performance and have quickly become the de-facto standard for semantic segmentation, with the added benefit that task-specific feature design is no longer necessary. However, a major downside of deep learning methods is that they are extremely data hungry, thus aggravating the perennial bottleneck of supervised classification, to obtain enough annotated training data. On the other hand, it has been observed that they are rather robust against noise in the training labels. This opens up the intriguing possibility to avoid annotating huge amounts of training data, and instead train the classifier from existing legacy data or crowd-sourced maps that can exhibit high levels of noise. The question addressed in this paper is: can training with large-scale publicly available labels replace a substantial part of the manual labeling effort and still achieve sufficient performance? Such data will inevitably contain a significant portion of errors, but in return virtually unlimited quantities of it are available in larger parts of the world. We adapt a state-of-the-art CNN architecture for semantic segmentation of buildings and roads in aerial images, and compare its performance when using different training data sets, ranging from manually labeled pixel-accurate ground truth of the same city to automatic training data derived from OpenStreetMap data from distant locations. We report our results that indicate that satisfying performance can be obtained with significantly less manual annotation effort, by exploiting noisy large-scale training data.