Iterative spatial leave-one-out cross-validation and gap-filling based data augmentation for supervised learning applications in marine remote sensing

Iterative spatial leave-one-out cross-validation and gap-filling based data augmentation for supervised learning applications in marine remote sensing
复制标题

DOI:
10.1080/15481603.2022.2107113
复制
发表时间:
2022-08
影响因子:
6.7
通讯作者:
Andy Stock;A. Subramaniam
Andy Stock;A. Subramaniam
中科院分区:
地球科学2区
文献类型:
--
作者:
Andy Stock;A. Subramaniam

文献摘要

被引文献

相似文献

在海洋遥感中,监督学习可以将海洋表面附近的原位测量变量与可以从空间测量的变量联系起来。然而,用于训练和验证此类经验卫星算法的原位数据通常在空间上自相关和聚类,从而产生各种统计挑战,例如对空间结构的过拟合。此外,由于从研究船收集数据的费用和频繁的云层覆盖,在海洋中很少同时进行现场和卫星测量。我们提出了两种方法来缓解这些挑战。第一种方法建立在空间留一交叉验证(slocv)的基础上,该方法旨在通过强制训练和测试观测值之间的最小分离距离,在数据空间自相关时提供可靠的误差估计。然而,对于稀疏和空间聚类的数据,估计这个距离可能是不可能的。因此,我们建议在一定距离范围内迭代和集成误差估计(iSLOOCV)。为了解决基于海洋原位数据的标记数据集通常规模较小的问题,我们测试了通过海洋卫星数据的云填充算法增加算法训练的观测数量是否可以改善预测。这两种方法的潜力是通过开发经验算法来绘制七种诊断色素(dp)的比例来证明的,这些色素可以作为墨西哥湾北部浮游植物群落组成的代理。我们使用不同的卫星数据产品集作为输入,利用iSLOOCV估计了13种算法的预测精度,并为7个dp中的4个找到了合适的算法。将海洋颜色和环境变量作为输入的随机森林总体上预测误差最低。iSLOOCV估计的预测值与观测值之间的相关性为0.69 ~ 0.85,平均绝对误差为0.02 ~ 0.13。这些DPs的每日地图和长期合成图与先前发表的结果大致一致。总的来说,当外推到更大的距离时,误差会增加,这突出了iSLOOCV如何说明基于子区域数据覆盖的算法性能变化。通过先前的空白填充生成更大的训练集,大大改善了7个DPs中3个DPs的所有误差测量,其他DPs的结果好坏参半。因此,通过卫星数据的空白填充来增加数据不应该被用作默认方法,但当怀疑监督学习应用受到训练集大小的限制时,它可能是一个有用的工具。
ABSTRACT In marine remote sensing, supervised learning can link variables measured in-situ near the ocean surface to variables that can be measured from space. However, the in-situ data used for training and validating such empirical satellite algorithms are often spatially auto-correlated and clustered, giving rise to various statistical challenges such as overfitting to spatial structures. Furthermore, co-located in-situ and satellite measurements are rare in the oceans because of the cost of data collection from research vessels and frequent cloud cover. We propose two methods to mitigate these challenges. The first method builds on spatial leave-one-out cross-validation (SLOOCV), an approach designed to provide sound error estimates when data are spatially auto-correlated by enforcing a minimum separation distance between training and test observations. However, estimating this distance may be impossible with sparse and spatially clustered data. We hence propose to iterate and integrate error estimates over a range of separation distances (iSLOOCV). To address the often-small size of labeled data sets based on marine in-situ data, we tested if increasing the number of observations for algorithm training by means of cloud-filling algorithms for marine satellite data improved predictions. The potential of these two methods is demonstrated by developing empirical algorithms for mapping the proportions of seven diagnostic pigments (DPs) that serve as proxies for phytoplankton community composition in the northern Gulf of Mexico. We estimated the prediction accuracy of 13 algorithms with iSLOOCV, using various sets of satellite data products as input, and found adequate algorithms for 4 of the 7 DPs. Random forests combining ocean color and environmental variables as input had the lowest prediction errors overall. Correlations between predictions and observations estimated by iSLOOCV ranged from 0.69 to 0.85 and mean absolute errors from 0.02 to 0.13. Daily maps and longer-term composites of these DPs were broadly consistent with previously published results. Overall, errors increased when extrapolating over larger distances, highlighting how iSLOOCV can illuminate changes in algorithm performance based on sub-regional data coverage. Generating larger training sets by prior gap-filling substantially improved all error measures for 3 of the 7 DPs, with mixed results for the others. Therefore, data augmentation by gap-filling of satellite data should not be used as a default approach but can be a useful tool when supervised learning applications are suspected to be limited by the size of the training set.