Multi-sensor cloud and cloud shadow segmentation with a convolutional neural network

Multi-sensor cloud and cloud shadow segmentation with a convolutional neural network
复制标题

DOI:
10.1016/j.rse.2019.05.022
复制
发表时间:
2019-09
影响因子:
13.5
通讯作者:
M. Wieland;Yu Li;S. Martinis
M. Wieland;Yu Li;S. Martinis
中科院分区:
工程技术1区
文献类型:
--
作者:
M. Wieland;Yu Li;S. Martinis

文献摘要

被引文献

相似文献

云和云阴影分割是任何使用多光谱卫星图像的应用的关键预处理步骤。特别是,与灾害有关的应用程序(例如,洪水监测或快速损害绘图),这些都是高度时间和数据关键的,需要在短时间内产生准确的云和云阴影掩模,同时能够适应目标域中的大变化(由大气条件、不同传感器、场景属性等引起)的方法。在这项研究中,我们提出了一种数据驱动的方法,基于修改后的U-Net卷积神经网络对单日期图像中的云和云阴影进行语义分割,旨在满足这些要求。我们训练网络的全球数据库Landsat OLI图像分割的五个类(“阴影”,“云”,“水”,“土地”和“雪/冰”)。我们将结果与最先进的方法进行比较,证明模型在多个卫星传感器(Landsat TM,Landsat ETM+,Landsat OLI和Sentinel-2)上的泛化能力,并显示不同训练策略和光谱波段组合对分割性能的影响。我们的方法始终优于Fmask和传统的随机森林分类器在全球分布的多传感器测试数据集的准确性,科恩的Kappa系数,骰子系数和推理速度。结果表明,一个减少的特征空间组成的红色,绿色,蓝色和近红外波段已经产生了良好的结果,所有测试的传感器。如果可能的话,增加短波红外波段可以提高准确性。训练数据的对比度和亮度增强进一步提高了分割性能。性能最好的U-Net模型实现了0.89的准确度,Kappa为0.82,Dice系数为0.85,同时以44.8 s/megapixel(GPU上为2.8 s/megapixel)对896个测试图像块进行推理。随机森林分类器在相同的训练和测试数据上达到0.79的准确度,Kappa为0.65,Dice系数为0.74,推理时间为3.9 s/megapixel(CPU上)。基于规则的Fmask方法需要更长的时间(277.8 s/megapixel),并产生精度为0.75,Kappa为0.60,Dice系数为0.72的结果。
Cloud and cloud shadow segmentation is a crucial pre-processing step for any application that uses multi-spectral satellite images. In particular, disaster related applications (e.g., flood monitoring or rapid damage mapping), which are highly time- and data-critical, require methods that produce accurate cloud and cloud shadow masks in short time while being able to adapt to large variations in the target domain (induced by atmospheric conditions, different sensors, scene properties, etc.). In this study, we propose a data-driven approach to semantic segmentation of cloud and cloud shadow in single date images based on a modified U-Net convolutional neural network that aims to fulfil these requirements. We train the network on a global database of Landsat OLI images for the segmentation of five classes (“shadow”, “cloud”, “water”, “land” and “snow/ice”). We compare the results to state-of-the-art methods, proof the model's generalization ability across multiple satellite sensors (Landsat TM, Landsat ETM+, Landsat OLI and Sentinel-2) and show the influence of different training strategies and spectral band combinations on the performance of the segmentation. Our method consistently outperforms Fmask and a traditional Random Forest classifier on a globally distributed multi-sensor test dataset in terms of accuracy, Cohen's Kappa coefficient, Dice coefficient and inference speed. The results indicate that a reduced feature space composed solely of red, green, blue and near-infrared bands already produces good results for all tested sensors. If available, adding shortwave-infrared bands can increase the accuracy. Contrast and brightness augmentations of the training data further improve the segmentation performance. The best performing U-Net model achieves an accuracy of 0.89, Kappa of 0.82 and Dice coefficient of 0.85, while running the inference over 896 test image tiles with 44.8 s/megapixel (2.8 s/megapixel on GPU). The Random Forest classifier reaches an accuracy of 0.79, Kappa of 0.65 and Dice coefficient of 0.74 with 3.9 s/megapixel inference time (on CPU) on the same training and testing data. The rule-based Fmask method takes significantly longer (277.8 s/megapixel) and produces results with an accuracy of 0.75, Kappa of 0.60 and Dice coefficient of 0.72.