High-Dimensional Bayesian Geostatistics.

High-Dimensional Bayesian Geostatistics.
复制标题

DOI:
10.1214/17-ba1056r
复制
发表时间:
2017-06
期刊:
影响因子:
4.4
通讯作者:
Banerjee S
Banerjee S
中科院分区:
数学2区
文献类型:
--
作者:
Banerjee S

文献摘要

被引文献

相似文献

随着地理信息系统 (GIS) 和用户友好型软件功能的不断增强,当今的统计学家经常会遇到包含大量空间位置和时间点观测结果的地理参考数据。在过去的十年中,分层时空过程模型已成为广泛部署的统计工具,供研究人员更好地理解时空变异性的复杂性质。然而,拟合分层时空模型通常涉及昂贵的矩阵计算,其复杂性随着空间位置和时间点的数量以三次方顺序增加。这使得此类模型对于大型数据集不可行。本文重点回顾了两种构建明确定义的高度可扩展的时空随机过程的方法。这两个过程都可以用作时空随机场的“先验”。第一种方法构建在低维子空间上运行的低秩过程。第二种方法构建一个最近邻高斯过程(NNGP),确保其有限实现的稀疏精度矩阵。这两个过程都可以用作嵌入在丰富的分层建模框架中的可扩展先验,以提供完整的贝叶斯推理。这些方法可以描述为针对大型时空数据集的基于模型的解决方案。该模型确保算法复杂性具有约 n 个浮点运算(触发器),其中 n 是空间位置的数量(每次迭代)。我们比较这些方法并提供对其方法基础的一些见解。
With the growing capabilities of Geographic Information Systems (GIS) and user-friendly software, statisticians today routinely encounter geographically referenced data containing observations from a large number of spatial locations and time points. Over the last decade, hierarchical spatiotemporal process models have become widely deployed statistical tools for researchers to better understand the complex nature of spatial and temporal variability. However, fitting hierarchical spatiotemporal models often involves expensive matrix computations with complexity increasing in cubic order for the number of spatial locations and temporal points. This renders such models unfeasible for large data sets. This article offers a focused review of two methods for constructing well-defined highly scalable spatiotemporal stochastic processes. Both these processes can be used as “priors” for spatiotemporal random fields. The first approach constructs a low-rank process operating on a lower-dimensional subspace. The second approach constructs a Nearest-Neighbor Gaussian Process (NNGP) that ensures sparse precision matrices for its finite realizations. Both processes can be exploited as a scalable prior embedded within a rich hierarchical modeling framework to deliver full Bayesian inference. These approaches can be described as model-based solutions for big spatiotemporal datasets. The models ensure that the algorithmic complexity has ~ n floating point operations (flops), where n the number of spatial locations (per iteration). We compare these methods and provide some insight into their methodological underpinnings.