SPARSE PRINCIPAL COMPONENT ANALYSIS AND ITERATIVE THRESHOLDING

SPARSE PRINCIPAL COMPONENT ANALYSIS AND ITERATIVE THRESHOLDING
复制标题

DOI:
10.1214/13-aos1097
复制
发表时间:
2013-04-01
影响因子:
4.5
通讯作者:
Ma, Zongming
Ma, Zongming
中科院分区:
数学1区
文献类型:
--
作者:
Ma, Zongming

文献摘要

被引文献

相似文献

主成分分析(PCA)是一种经典的降维方法,它将数据投影到由协方差矩阵的前导特征向量构成的主子空间上。然而,当特征的数量p与样本大小n相当或甚至比样本大小n大得多时,它表现得很差。在本文中,我们提出了一个新的迭代阈值方法估计主子空间的设置中的领导特征向量稀疏。在尖峰协方差模型下,我们发现新方法在一系列高维稀疏设置中一致地恢复了主子空间和前导特征向量,甚至是最优的。仿真算例也证明了该算法的优越性。
Principal component analysis (PCA) is a classical dimension reduction method which projects data onto the principal subspace spanned by the leading eigenvectors of the covariance matrix. However, it behaves poorly when the number of features p is comparable to, or even much larger than, the sample size n. In this paper, we propose a new iterative thresholding approach for estimating principal subspaces in the setting where the leading eigenvectors are sparse. Under a spiked covariance model, we find that the new approach recovers the principal subspace and leading eigenvectors consistently, and even optimally, in a range of high-dimensional sparse settings. Simulated examples also demonstrate its competitive performance.