The Masked Sample Covariance Estimator: An Analysis via Matrix Concentration Inequalities

The Masked Sample Covariance Estimator: An Analysis via Matrix Concentration Inequalities
复制标题

掩蔽样本协方差估计器:通过矩阵浓度不等式进行分析

DOI:
10.1093/imaiai/ias001
复制
发表时间:
2011
期刊:
Information and Inference: A Journal of the IMA
影响因子:
--
通讯作者:
J. Tropp
J. Tropp
中科院分区:
--
文献类型:
--
作者:
Richard Y. Chen;Alex Gittens;J. Tropp

文献摘要

被引文献

相似文献

在变量数量p超过可用于构建估计的样本数量n的情况下,协方差估计变得具有挑战性。避免这个问题的一种方法是假设协方差矩阵几乎是稀疏的,并且只专注于估计有意义的条目。为了分析这种方法,Levina和Vershynin(2011)引入了一种称为掩蔽协方差估计的形式主义,其中样本协方差估计的每个条目都被重新加权,以反映对其重要性的先验评估。本文利用一个矩阵浓缩不等式对掩蔽样本协方差估计进行了简短的分析。主要结果适用于具有至少四个矩的一般分布。该理论专门针对高斯分布的情况,比以前的工作提供了质量上的改进。例如,新的结果表明,与Levina和Vershynin得到的样本复杂性n=O(Blog^5p)相比,n=O(Blog^2p)样本足以估计带宽B高达相对谱范数误差的带状协方差矩阵。
Covariance estimation becomes challenging in the regime where the number p of variables outstrips the number n of samples available to construct the estimate. One way to circumvent this problem is to assume that the covariance matrix is nearly sparse and to focus on estimating only the significant entries. To analyze this approach, Levina and Vershynin (2011) introduce a formalism called masked covariance estimation, where each entry of the sample covariance estimator is reweighted to reflect an a priori assessment of its importance. This paper provides a short analysis of the masked sample covariance estimator by means of a matrix concentration inequality. The main result applies to general distributions with at least four moments. Specialized to the case of a Gaussian distribution, the theory offers qualitative improvements over earlier work. For example, the new results show that n = O(B log^2 p) samples suffice to estimate a banded covariance matrix with bandwidth B up to a relative spectral-norm error, in contrast to the sample complexity n = O(B log^5 p) obtained by Levina and Vershynin.