Accuracy of Pseudo-Inverse Covariance Learning-A Random Matrix Theory Analysis

Accuracy of Pseudo-Inverse Covariance Learning-A Random Matrix Theory Analysis
复制标题

DOI:
10.1109/tpami.2010.186
复制
发表时间:
2011-07-01
影响因子:
23.6
通讯作者:
Hoyle, David C.
Hoyle, David C.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hoyle, David C.

文献摘要

被引文献

相似文献

对于许多学习问题,需要估计逆总体协方差,并且通常通过反转样本协方差矩阵来获得。对于现代科学数据集,样本点的数量越来越少,因此样本协方差不可逆。在这种情况下,由对应于非零样本协方差特征值的特征向量构造的Moore-Penrose伪逆样本协方差矩阵通常用作逆总体协方差矩阵的近似。伪逆样本协方差矩阵在估计真逆协方差时的重构误差可以通过两者之差的Frobenius范数来量化。重建误差主要由最小的非零样本协方差特征值和发散的样本大小变得相当的功能。对于高维数据,我们使用随机矩阵理论的技术和结果来研究一类广泛的人口协方差矩阵的重建误差。我们还展示了如何装袋和随机子空间方法可以减少重建误差,并可以结合起来,以提高利用伪逆样本协方差矩阵的分类器的准确性。我们测试我们的分析模拟和基准数据集。
For many learning problems, estimates of the inverse population covariance are required and often obtained by inverting the sample covariance matrix. Increasingly for modern scientific data sets, the number of sample points is less than the number of features and so the sample covariance is not invertible. In such circumstances, the Moore-Penrose pseudo-inverse sample covariance matrix, constructed from the eigenvectors corresponding to nonzero sample covariance eigenvalues, is often used as an approximation to the inverse population covariance matrix. The reconstruction error of the pseudo-inverse sample covariance matrix in estimating the true inverse covariance can be quantified via the Frobenius norm of the difference between the two. The reconstruction error is dominated by the smallest nonzero sample covariance eigenvalues and diverges as the sample size becomes comparable to the number of features. For high-dimensional data, we use random matrix theory techniques and results to study the reconstruction error for a wide class of population covariance matrices. We also show how bagging and random subspace methods can result in a reduction in the reconstruction error and can be combined to improve the accuracy of classifiers that utilize the pseudo-inverse sample covariance matrix. We test our analysis on both simulated and benchmark data sets.