Accuracy of Empirical Satellite Algorithms for Mapping Phytoplankton Diagnostic Pigments in the Open Ocean: A Supervised Learning Perspective

Accuracy of Empirical Satellite Algorithms for Mapping Phytoplankton Diagnostic Pigments in the Open Ocean: A Supervised Learning Perspective
复制标题

DOI:
10.3389/fmars.2020.00599
复制
发表时间:
2020-07
期刊:
Remote. Sens.
影响因子:
--
通讯作者:
A. Stock;A. Subramaniam
A. Stock;A. Subramaniam
中科院分区:
其他
文献类型:
--
作者:
A. Stock;A. Subramaniam

文献摘要

被引文献

相似文献

从空间监测浮游植物群落组成是海洋遥感的一个重要挑战。研究人员为此提出了几种算法。然而,用于在全球范围内训练和验证这种算法的现场数据通常是沿着沿着船只巡航轨迹和在一些经过充分研究的地点聚集的,而许多大型海洋区域根本没有现场数据。此外,海洋学变量通常在空间上是自相关的。在这种情况下,用随机选择的观测值验证算法的常见做法可能会低估误差。基于全球原位HPLC数据库,我们应用监督学习方法来训练和测试经验算法,预测作为不同浮游植物类型生物标志物的8种诊断色素的相对浓度。对于每种色素,我们训练了三种类型的卫星算法,这些算法由其输入数据区分开来:基于丰度的算法(仅使用叶绿素a作为输入)、光谱算法(使用遥感反射率)和生态算法(结合反射率和环境变量)。这些算法被实现为统计模型(平滑样条、多项式、随机森林和提升回归树)。为了解决数据的聚类和空间自相关性,我们通过空间块交叉验证来测试算法。这提供了一个不太有信心的图片的潜力,全球诊断色素,因此,相关的浮游植物类型使用现有的卫星数据比以前的一些研究和五倍交叉验证进行比较所建议的。在八种诊断色素中,只有两种(岩藻黄质和玉米黄质)可以在海洋区域预测,这些算法没有经过训练,误差比常数空模型低得多。因此,根据现有的多光谱卫星数据和通常可获得的环境变量,全球范围的算法可以估计相对的诊断色素浓度,从而区分某些大类的浮游植物类型,但对某些类别和某些海洋区域可能不准确。总体而言,生态算法的预测误差最低,这表明环境变量包含的信息的全球空间分布的浮游植物群体,没有捕获在多光谱遥感反射率和卫星派生的叶绿素a浓度。与空间聚类程度成反比的加权训练观察改善了预测。最后,我们的研究结果表明,更多的讨论的最佳方法,训练和验证经验卫星算法是必要的,如果在原位数据分布不均匀的研究区域和空间集群。
Monitoring phytoplankton community composition from space is an important challenge in ocean remote sensing. Researchers have proposed several algorithms for this purpose. However, the in situ data used to train and validate such algorithms at the global scale are often clustered along ship cruise tracks and in some well-studied locations, whereas many large marine regions have no in situ data at all. Furthermore, oceanographic variables are typically spatially auto-correlated. In this situation, the common practice of validating algorithms with randomly chosen held-out observations can underestimate errors. Based on a global database of in situ HPLC data, we applied supervised learning methods to train and test empirical algorithms predicting the relative concentrations of eight diagnostic pigments that serve as biomarkers for different phytoplankton types. For each pigment, we trained three types of satellite algorithms distinguished by their input data: abundance-based (using only chlorophyll-a as input), spectral (using remote sensing reflectance), and ecological algorithms (combining reflectance and environmental variables). The algorithms were implemented as statistical models (smoothing splines, polynomials, random forests, and boosted regression trees). To address clustering of data and spatial auto-correlation, we tested the algorithms by means of spatial block cross-validation. This provided a less confident picture of the potential for global mapping of diagnostic pigments and hence the associated phytoplankton types using existing satellite data than suggested by some previous research and a fivefold cross-validation conducted for comparison. Of the eight diagnostic pigments, only two (fucoxanthin and zeaxanthin) could be predicted in marine regions that the algorithms were not trained in with considerably lower errors than a constant null model. Thus, global-scale algorithms based on existing, multi-spectral satellite data and commonly available environmental variables can estimate relative diagnostic pigment concentrations and hence distinguish phytoplankton types in some broad classes, but are likely inaccurate for some classes and in some marine regions. Overall, the ecological algorithms had the lowest prediction errors, suggesting that environmental variables contain information about the global spatial distribution of phytoplankton groups that is not captured in multi-spectral remote sensing reflectance and satellite-derived Chl a concentrations. Weighting training observations inversely to the degree of spatial clustering improved predictions. Finally, our results suggest that more discussion of the best approaches for training and validating empirical satellite algorithms is needed if the in situ data are unevenly distributed in the study region and spatially clustered.