Covariance Estimation: The GLM and Regularization Perspectives

Covariance Estimation: The GLM and Regularization Perspectives
复制标题

DOI:
10.1214/11-sts358
复制
发表时间:
2011-08-01
影响因子:
5.7
通讯作者:
Pourahmadi, Mohsen
Pourahmadi, Mohsen
中科院分区:
数学2区
文献类型:
--
作者:
Pourahmadi, Mohsen

文献摘要

被引文献

相似文献

找到一个无约束的和统计上可解释的协方差矩阵的重新参数化仍然是一个开放的问题,在统计。它的解决方案是至关重要的协方差估计,特别是在最近的高维数据环境中,强制执行的正定性约束可能是计算昂贵的。我们从两个相对互补的角度对协方差矩阵建模所取得的进展进行了调查:(1)广义线性模型(GLM)或简约性和低维协变量的使用,以及(2)高维数据的正则化或稀疏性。一个新兴的,统一的和强大的趋势,在这两个方面是减少协方差估计问题,估计一系列回归问题。我们指出几个实例的回归为基础的配方。一个值得注意的情况是在稀疏估计的精度矩阵或高斯图形模型,导致快速图形LASSO算法。一些基于回归的Cholesky分解相对于经典的谱(特征值)和方差相关分解的优点和局限性突出。前者提供了一个无约束和统计上可解释的重新参数化,并保证了估计的协方差矩阵的正定性。它减少了不直观的任务,协方差估计的回归序列建模的成本强加一个先验秩序的变量。样本协方差矩阵的元素正则化,如条带化,锥形化和阈值化具有理想的渐近性质,并且稀疏估计的协方差矩阵是正定的,对于大样本和大维度,概率趋于1。
Finding an unconstrained and statistically interpretable reparameterization of a covariance matrix is still an open problem in statistics. Its solution is of central importance in covariance estimation, particularly in the recent high-dimensional data environment where enforcing the positive-definiteness constraint could be computationally expensive. We provide a survey of the progress made in modeling covariance matrices from two relatively complementary perspectives: (1) generalized linear models (GLM) or parsimony and use of covariates in low dimensions, and (2) regularization or sparsity for high-dimensional data. An emerging, unifying and powerful trend in both perspectives is that of reducing a covariance estimation problem to that of estimating a sequence of regression problems. We point out several instances of the regression-based formulation. A notable case is in sparse estimation of a precision matrix or a Gaussian graphical model leading to the fast graphical LASSO algorithm. Some advantages and limitations of the regression-based Cholesky decomposition relative to the classical spectral (eigenvalue) and variance-correlation decompositions are highlighted. The former provides an unconstrained and statistically interpretable reparameterization, and guarantees the positive-definiteness of the estimated covariance matrix. It reduces the unintuitive task of covariance estimation to that of modeling a sequence of regressions at the cost of imposing an a priori order among the variables. Elementwise regularization of the sample covariance matrix such as banding, tapering and thresholding has desirable asymptotic properties and the sparse estimated covariance matrix is positive definite with probability tending to one for large samples and dimensions.