Sparse estimation of a covariance matrix

Sparse estimation of a covariance matrix
复制标题

DOI:
10.1093/biomet/asr054
复制
发表时间:
2011-12-01
期刊:
影响因子:
2.7
通讯作者:
Tibshirani, Robert J.
Tibshirani, Robert J.
中科院分区:
数学2区
文献类型:
--
作者:
Bien, Jacob;Tibshirani, Robert J.

文献摘要

被引文献

相似文献

我们提出了一种方法,估计协方差矩阵的基础上,从一个多元正态分布的向量样本。特别地,我们在协方差矩阵的条目上使用套索惩罚来惩罚可能性。这种惩罚起着两个重要的作用:它减少了有效的参数数量,这是重要的,即使当向量的维数小于样本大小,因为参数的数量在变量的数量中二次增长,它产生一个估计是稀疏的。相反,稀疏逆协方差估计,我们的方法的近亲,这里达到的稀疏性是在协方差矩阵本身,而不是在逆矩阵。协方差矩阵中的零对应于边际独立性;因此,我们的方法在提供协方差的正定估计的同时执行模型选择。建议的惩罚最大似然问题是不凸的,所以我们使用一个优化最小化的方法,我们迭代求解凸近似原来的非凸问题。我们讨论了调整参数的选择,并在流式细胞仪数据集上演示了我们的方法如何产生变量之间关系的可解释的图形显示。我们进行模拟表明,简单的经验协方差矩阵的元素级阈值与我们的方法识别的稀疏结构是有竞争力的。此外,我们展示了我们的方法可以用来解决一个先前研究的特殊情况下,所需的稀疏模式是预先指定的。
We suggest a method for estimating a covariance matrix on the basis of a sample of vectors drawn from a multivariate normal distribution. In particular, we penalize the likelihood with a lasso penalty on the entries of the covariance matrix. This penalty plays two important roles: it reduces the effective number of parameters, which is important even when the dimension of the vectors is smaller than the sample size since the number of parameters grows quadratically in the number of variables, and it produces an estimate which is sparse. In contrast to sparse inverse covariance estimation, our method's close relative, the sparsity attained here is in the covariance matrix itself rather than in the inverse matrix. Zeros in the covariance matrix correspond to marginal independencies; thus, our method performs model selection while providing a positive definite estimate of the covariance. The proposed penalized maximum likelihood problem is not convex, so we use a majorize-minimize approach in which we iteratively solve convex approximations to the original nonconvex problem. We discuss tuning parameter selection and demonstrate on a flow-cytometry dataset how our method produces an interpretable graphical display of the relationship between variables. We perform simulations that suggest that simple elementwise thresholding of the empirical covariance matrix is competitive with our method for identifying the sparsity structure. Additionally, we show how our method can be used to solve a previously studied special case in which a desired sparsity pattern is prespecified.