On nearest-neighbor Gaussian process models for massive spatial data.

On nearest-neighbor Gaussian process models for massive spatial data.
复制标题

DOI:
10.1002/wics.1383
复制
发表时间:
2016-09
期刊:
Wiley interdisciplinary reviews. Computational statistics
影响因子:
--
通讯作者:
Gelfand AE
Gelfand AE
中科院分区:
其他
文献类型:
--
作者:
Datta A;Banerjee S;Finley AO;Gelfand AE

文献摘要

被引文献

相似文献

高斯过程(GP)模型为对位置和时间索引的数据集进行建模提供了一种非常灵活的非参数方法。然而,对于大型空间数据集,GP模型的存储和计算要求是不可行的。最近邻高斯过程(Datta A,Banerjee S,Finley AO,Gelfand AE。用于大型地统计数据集的分层最近邻高斯过程模型。《美国统计协会杂志》2016年,JASA)通过使用来自少数最近邻的局部信息提供了一种可扩展的替代方法。通过在模型的条件设定中使用邻域集来实现可扩展性。我们展示了这如何等同于对大型协方差矩阵的乔列斯基因子进行稀疏建模。我们还讨论了一种使用稀疏局部克里金法构建可扩展高斯过程的通用方法。我们进行了一项多元数据分析,该分析表明最近邻方法如何在速度快数倍的情况下得出与满秩GP无法区分的推断。最后,我们还提出了一种NNGP模型的变体,用于自动选择邻域集大小。
Gaussian Process (GP) models provide a very flexible nonparametric approach to modeling location-and-time indexed datasets. However, the storage and computational requirements for GP models are infeasible for large spatial datasets. Nearest Neighbor Gaussian Processes (Datta A, Banerjee S, Finley AO, Gelfand AE. Hierarchical nearest-neighbor gaussian process models for large geostatistical datasets. J Am Stat Assoc 2016., JASA) provide a scalable alternative by using local information from few nearest neighbors. Scalability is achieved by using the neighbor sets in a conditional specification of the model. We show how this is equivalent to sparse modeling of Cholesky factors of large covariance matrices. We also discuss a general approach to construct scalable Gaussian Processes using sparse local kriging. We present a multivariate data analysis which demonstrates how the nearest neighbor approach yields inference indistinguishable from the full rank GP despite being several times faster. Finally, we also propose a variant of the NNGP model for automating the selection of the neighbor set size.