Geostatistical interpolation of positively skewed and censored data in a dioxin-contaminated site

Geostatistical interpolation of positively skewed and censored data in a dioxin-contaminated site
复制标题

DOI:
10.1021/es991450y
复制
发表时间:
2000-10-01
影响因子:
11.4
通讯作者:
Goovaerts, P
Goovaerts, P
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Saito, H;Goovaerts, P

文献摘要

被引文献

相似文献

污染场地危险区域的正确划分首先依赖于污染物浓度的准确预测,这项任务通常因审查数据的存在而变得复杂(低于检测限的观测值和高度正偏斜的直方图。本文比较了四种地质统计算法(普通克里金法、对数正态克里金法、多重高斯克里金法和指示克里金法)的预测性能 通过一组 600 个二恶英浓度的交叉验证。尽管存在理论上的局限性,对数正态克里金法始终能产生最佳结果(最小的预测误差、最少的误报和最低的总成本)。对于一系列采样强度,交叉验证已重复 100 次,这降低了这些结果仅反映采样波动、指示克里金法 (IK) 的风险,在简化实施中 中位 IK 可以产生良好的预测,但由于低估高二恶英浓度而导致中度偏差。普通克里金法受数据稀疏性的影响最大,导致很大一部分修复单位在使用少于 100 个观测值时错误地宣布受到污染。最后,基于多高斯克里金估计的决策成本最高,并且会产生很大比例的误报,而这些误报无法通过收集来减少 额外的样本。
A correct delineation of hazardous areas in a contaminated site relies first on accurate predictions of pollutant concentrations, a task usually complicated by the presence of censored data (observations below the detection limit and highly positively skewed histograms. This paper compares the prediction performances of four geostatistical algorithms (ordinary kriging, log-normal kriging, multi-Gaussian kriging, and indicator kriging) through the cross validation of a set of 600 dioxin concentrations. Despite its theoretical limitations, log-normal kriging consistently yields the best results (smallest prediction errors, least false positives, and lowest total costs). The cross validation has been repeated 100 times for a series of sampling intensities, which reduces the risk that these results simply reflect sampling fluctuations, indicator kriging (IK), in the simplified implementation of median IK, produces good predictions except for a moderate bias caused by the underestimation of high dioxin concentrations. Ordinary kriging is the most affected by data sparsity, leading to a large proportion of remediation units wrongly declared contaminated when less than 100 observations were used. Last, decisions based on multi-Gaussian kriging estimates are the most costly and create a large proportion of false positives that cannot be reduced by the collection of additional samples.