Quantification of Soil Variables in a Heterogeneous Soil Region With VIS–NIR–SWIR Data Using Different Statistical Sampling and Modeling Strategies

Quantification of Soil Variables in a Heterogeneous Soil Region With VIS–NIR–SWIR Data Using Different Statistical Sampling and Modeling Strategies
复制标题

DOI:
10.1109/jstars.2016.2572879
复制
发表时间:
2016-07
影响因子:
5.5
通讯作者:
Michael Vohland;M. Harbich;M. Ludwig;C. Emmerling;S. Thiele-Bruhn
Michael Vohland;M. Harbich;M. Ludwig;C. Emmerling;S. Thiele-Bruhn
中科院分区:
工程技术3区
文献类型:
--
作者:
Michael Vohland;M. Harbich;M. Ludwig;C. Emmerling;S. Thiele-Bruhn

文献摘要

被引文献

相似文献

从分光辐射计数据获得的土壤性质的估计精度显着依赖于个人的样本集。对校准集进行采样的统计方法的选择以及具有装袋和/或光谱变量选择的多变量建模方法的扩展可以优化预测。我们研究了一组172耕地表土从特里尔(德国)附近的一个地区,覆盖-通常是典型的中到大规模应用的土壤光谱-广泛的不同的土壤情况。然而,关于目标变量-有机碳(OC),氮(N),微生物生物量(Cmic)和热稳定碳(Cinert)-的差异很小。基于分割的校准和验证数据与Kennard-Stone算法,我们发现只有适度的改进,对偏最小二乘回归(PLSR)相结合时,PLSR与装袋,光谱变量的选择,与“竞争性自适应加权采样”(汽车)。在验证中,OC(从0.75到0.79)、N(从0.72到0.77)和Cinert(从0.66到0.68)的R2有所改善。此外,我们为每个验证样品使用了单独的校准集。在这种“局部”方法中,我们在光谱特征空间中对校准样本进行聚类,并从每个聚类中单独选择最相似的样本。将装袋-CARS-PLSR与这种局部方法相结合,将Cinert的R2显著提高到0.76,OC的R2略微提高到0.82,Cmic的R2提高到0.76(之前为0.73)。局部方法的效果是双重的,因为它从校准中去除了不适当的样本,并平衡了数据分布中的偏度。
Estimation accuracies obtained for soil properties from spectroradiometer data markedly depend on the individual sample set. The choice of the statistical method to sample a calibration set and the extension of the multivariate modeling approach with bagging and/or spectral variable selection may optimize predictions. We studied this with a set of 172 arable topsoils from a region near Trier (Germany) that covered-as often typical for medium to large-scale applications of soil spectroscopy-a wide range of different soil situations. Yet, differences concerning target variables-organic carbon (OC), nitrogen (N), microbial biomass (Cmic) and thermostable carbon (Cinert)-were small. Based on a split of calibration and validation data with the Kennard-Stone algorithm, we found only moderate improvements towards partial least squares regression (PLSR) when combining PLSR with bagging and, for spectral variable selection, with “competitive adaptive reweighted sampling” (CARS). R2 improved for OC (from 0.75 to 0.79), N (from 0.72 to 0.77) and Cinert (from 0.66 to 0.68) in the validation. Additionally, we used individual calibration sets for each validation sample. In this “local” approach, we clustered calibration samples in the spectral feature space and selected individually the most similar sample from each cluster. Combining bagging-CARS-PLSR with this local approach improved R2 markedly to 0.76 for Cinert, and slightly to 0.82 for OC and to 0.76 (previously 0.73) for Cmic. Effects of the local approach were twofold, as it removed improper samples from the calibration and balanced skewness in the data distribution.