Gaussian predictive process models for large spatial data sets.

Gaussian predictive process models for large spatial data sets.
复制标题

DOI:
10.1111/j.1467-9868.2008.00663.x
复制
发表时间:
2008-09-01
期刊:
Journal of the Royal Statistical Society. Series B, Statistical methodology
影响因子:
--
通讯作者:
Sang H
Sang H
中科院分区:
其他
文献类型:
--
作者:
Banerjee S;Gelfand AE;Finley AO;Sang H

文献摘要

被引文献

相似文献

有了地理编码位置的科学数据,研究人员越来越多地转向空间过程模型来进行统计推断。在过去的十年中,通过马尔可夫链蒙特卡罗方法实现的分层模型在空间建模中变得特别流行,因为它们具有灵活性和能力,可以拟合经典方法无法实现的模型,并且可以避免可能不适当的渐近。然而,拟合层次空间模型往往涉及昂贵的矩阵分解,其计算复杂度随着空间位置的数量呈三次增长,使得这种模型不适合大型空间数据集。这种计算负担加剧了多变量设置与几个空间相关的响应变量。在频繁的时间点收集数据和使用时空过程模型时,这一问题也会加剧。对于这一挑战,我们的贡献是使用我们称之为空间和时空数据的预测过程模型。每一个空间(或时空)过程都会诱发一个预测过程模型(实际上是任意多个)。后者将前者的过程实现投射到较低维的子空间,从而减少了计算负担。因此,我们实现了在大数据集背景下适应非平稳、非高斯、可能多变量、可能时空过程的灵活性。我们讨论了这些预测过程的吸引人的理论性质。我们还提供了一个包含这些不同设置的计算模板。最后,我们用模拟和真实数据集来说明该方法。
With scientific data available at geocoded locations, investigators are increasingly turning to spatial process models for carrying out statistical inference. Over the last decade, hierarchical models implemented through Markov chain Monte Carlo methods have become especially popular for spatial modelling, given their flexibility and power to fit models that would be infeasible with classical methods as well as their avoidance of possibly inappropriate asymptotics. However, fitting hierarchical spatial models often involves expensive matrix decompositions whose computational complexity increases in cubic order with the number of spatial locations, rendering such models infeasible for large spatial data sets. This computational burden is exacerbated in multivariate settings with several spatially dependent response variables. It is also aggravated when data are collected at frequent time points and spatiotemporal process models are used. With regard to this challenge, our contribution is to work with what we call predictive process models for spatial and spatiotemporal data. Every spatial (or spatiotemporal) process induces a predictive process model (in fact, arbitrarily many of them). The latter models project process realizations of the former to a lower dimensional subspace, thereby reducing the computational burden. Hence, we achieve the flexibility to accommodate non-stationary, non-Gaussian, possibly multivariate, possibly spatiotemporal processes in the context of large data sets. We discuss attractive theoretical properties of these predictive processes. We also provide a computational template encompassing these diverse settings. Finally, we illustrate the approach with simulated and real data sets.