Sparse Precision Matrix Selection for Fitting Gaussian Random Field Models to Large Data Sets

Sparse Precision Matrix Selection for Fitting Gaussian Random Field Models to Large Data Sets
复制标题

用于将高斯随机场模型拟合到大型数据集的稀疏精度矩阵选择

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
E. Castillo
E. Castillo
中科院分区:
--
文献类型:
--
作者:
S. Tajbakhsh;N. Aybat;E. Castillo

文献摘要

被引文献

相似文献

通过最大似然(ML)将高斯随机场(GRF)模型拟合到空间数据的迭代方法每次迭代需要$mathcal{O}(n^3)$浮点运算,其中$n$表示数据位置的数量。对于大型数据集,每次迭代的$mathcal{O}(n^3)$复杂度以及ML问题的非凸性使得传统ML方法对于GRF拟合效率低下。对于各向异性GRF,问题甚至更加严重,其中协方差函数参数的数量随着过程域维度而增加。在本文中,我们提出了一个新的两步GRF估计过程时,是二阶平稳。首先,利用观测点之间的距离信息,利用乘子交替方向法,求解一个用加权$ell_1$-范数正则化的凸似然问题,以拟合一个稀疏的协方差矩阵.其次,通过求解最小二乘问题估计GRF空间协方差函数的参数。所提出的估计的理论误差界提供,此外,收敛的估计显示为每个位置的样本数量的增加。所提出的方法进行了数值比较与国家的最先进的方法为大$n$。数据分割方案的实施,以处理大型数据集。
Iterative methods for fitting a Gaussian Random Field (GRF) model to spatial data via maximum likelihood (ML) require $mathcal{O}(n^3)$ floating point operations per iteration, where $n$ denotes the number of data locations. For large data sets, the $mathcal{O}(n^3)$ complexity per iteration together with the non-convexity of the ML problem render traditional ML methods inefficient for GRF fitting. The problem is even more aggravated for anisotropic GRFs where the number of covariance function parameters increases with the process domain dimension. In this paper, we propose a new two-step GRF estimation procedure when the process is second-order stationary. First, a emph{convex} likelihood problem regularized with a weighted $ell_1$-norm, utilizing the available distance information between observation locations, is solved to fit a sparse emph{{precision} (inverse covariance) matrix to the observed data using the Alternating Direction Method of Multipliers. Second, the parameters of the GRF spatial covariance function are estimated by solving a least squares problem. Theoretical error bounds for the proposed estimator are provided; moreover, convergence of the estimator is shown as the number of samples per location increases. The proposed method is numerically compared with state-of-the-art methods for big $n$. Data segmentation schemes are implemented to handle large data sets.