A Case Study Competition Among Methods for Analyzing Large Spatial Data

A Case Study Competition Among Methods for Analyzing Large Spatial Data
复制标题

DOI:
10.1007/s13253-018-00348-w
复制
发表时间:
2019-09-01
影响因子:
1.4
通讯作者:
Zammit-Mangion, Andrew
Zammit-Mangion, Andrew
中科院分区:
数学4区
文献类型:
--
作者:
Heaton, Matthew J.;Datta, Abhirup;Zammit-Mangion, Andrew

文献摘要

被引文献

相似文献

高斯过程是空间数据分析中不可缺少的工具。然而,随着“大数据”时代的到来,传统的高斯过程在现代空间数据计算上已经不可行。因此,人们提出了各种替代全高斯过程的方法,这些方法更适合处理大空间数据。这些现代方法通常利用低阶结构和/或多核和多线程计算环境来促进计算。本研究首先介绍了分析大型空间数据的几种方法。其次,本研究描述了所描述的方法之间的预测竞争的结果,这些方法是由不同的团队在方法上具有很强的专业知识。具体来说,每个研究小组都提供了两个训练数据集(一个模拟数据集,一个观察数据集)以及一组预测位置。然后,每个小组都编写了自己的方法实现,以在给定的位置产生预测,然后每个方法都在公共计算环境中运行。然后在各种预测诊断方面比较了这些方法。关于方法和代码的实现细节的补充材料可以在线获取本文。
The Gaussian process is an indispensable tool for spatial data analysts. The onset of the "big data" era, however, has lead to the traditional Gaussian process being computationally infeasible for modern spatial data. As such, various alternatives to the full Gaussian process that are more amenable to handling big spatial data have been proposed. These modern methods often exploit low-rank structures and/or multi-core and multi-threaded computing environments to facilitate computation. This study provides, first, an introductory overview of several methods for analyzing large spatial data. Second, this study describes the results of a predictive competition among the described methods as implemented by different groups with strong expertise in the methodology. Specifically, each research group was provided with two training datasets (one simulated and one observed) along with a set of prediction locations. Each group then wrote their own implementation of their method to produce predictions at the given location and each was subsequently run on a common computing environment. The methods were then compared in terms of various predictive diagnostics. Supplementary materials regarding implementation details of the methods and code are available for this article online.