ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data

ResUNet-a: A deep learning framework for semantic segmentation of remotely sensed data
复制标题

DOI:
10.1016/j.isprsjprs.2020.01.013
复制
发表时间:
2020-04-01
影响因子:
12.7
通讯作者:
Wu, Chen
Wu, Chen
中科院分区:
工程技术1区
文献类型:
--
作者:
Diakogiannis, Foivos, I;Waldner, Francois;Wu, Chen

文献摘要

被引文献

相似文献

高分辨率航空影像的场景理解对于各种遥感应用中的自动化监测任务具有重要意义。由于感兴趣的对象的像素值的类内方差大,类间方差小,这仍然是一项具有挑战性的任务。近年来,深度卷积神经网络已开始用于遥感应用,并展示了对象像素级分类的最新性能。在这里,我们提出了一个可靠的框架,为任务的语义分割的monotemporal非常高分辨率的航空图像的performant结果。我们的框架由一个新的深度学习架构ResUNet-a和一个基于Dice损失的新损失函数组成。ResUNet-a使用UNet编码器/解码器骨干,结合残差连接,atrous卷积,金字塔场景解析池和多任务推理。ResUNet-a依次推断对象的边界、分割掩码的距离变换、分割掩码和输入的彩色重建。每个任务都以前一个任务的推理为条件,从而在各个任务之间建立条件关系,这是通过架构的计算图来描述的。我们分析了几种口味的广义骰子损失的语义分割的性能,我们引入了一种新的变体损失函数的对象,具有良好的收敛性能,即使在存在高度不平衡的类的语义分割。我们的建模框架的性能进行评估的ISPRS二维波茨坦数据集。结果显示了最先进的性能,对于我们的最佳模型,所有类别的平均F1得分为92.9%。
Scene understanding of high resolution aerial images is of great importance for the task of automated monitoring in various remote sensing applications. Due to the large within-class and small between-class variance in pixel values of objects of interest, this remains a challenging task. In recent years, deep convolutional neural networks have started being used in remote sensing applications and demonstrate state of the art performance for pixel level classification of objects. Here we propose a reliable framework for performant results for the task of semantic segmentation of monotemporal very high resolution aerial images. Our framework consists of a novel deep learning architecture, ResUNet-a, and a novel loss function based on the Dice loss. ResUNet-a uses a UNet encoder/decoder backbone, in combination with residual connections, atrous convolutions, pyramid scene parsing pooling and multi-tasking inference. ResUNet-a infers sequentially the boundary of the objects, the distance transform of the segmentation mask, the segmentation mask and a colored reconstruction of the input. Each of the tasks is conditioned on the inference of the previous ones, thus establishing a conditioned relationship between the various tasks, as this is described through the architecture's computation graph. We analyse the performance of several flavours of the Generalized Dice loss for semantic segmentation, and we introduce a novel variant loss function for semantic segmentation of objects that has excellent convergence properties and behaves well even under the presence of highly imbalanced classes. The performance of our modeling framework is evaluated on the ISPRS 2D Potsdam dataset. Results show state-of-the-art performance with an average F1 score of 92.9% over all classes for our best model.