A New Basis for Sparse Principal Component Analysis

A New Basis for Sparse Principal Component Analysis
复制标题

稀疏主成分分析的一种新基础

DOI:
10.1080/10618600.2023.2256502
复制
发表时间:
2020-07
影响因子:
2.4
通讯作者:
Fan Chen;Karl Rohe
Fan Chen;Karl Rohe
中科院分区:
数学2区
文献类型:
--
作者:
Fan Chen;Karl Rohe

文献摘要

相似文献

摘要稀疏主成分分析(PCA)以前的版本都假定特征基(p × k矩阵)是近似稀疏的。我们提出了一种方法,假设p × k矩阵在k × k旋转后变得近似稀疏。该算法的最简单版本是使用前导k个主成分进行排序。然后,对主成分进行k × k正交旋转,使其近似稀疏。最后,对旋转后的主成分进行软阈值处理。这种方法与现有方法不同,因为它使用正交旋转来近似稀疏基。一个结果是稀疏分量不需要是前导特征向量,而是它们的混合。通过这种方式,我们提出了一个新的(旋转)稀疏PCA的基础。此外,我们的方法避免了“通货紧缩”和多个调优参数所需的。我们的稀疏PCA框架是通用的;例如,它自然地扩展到数据矩阵的双向分析,以同时降低行和列的维度。我们提供的证据表明,对于相同水平的稀疏性,所提出的稀疏PCA方法更稳定,可以解释更多的方差相比,替代方法。通过三个应用程序-稀疏编码的图像,转录组测序数据的分析,和大规模聚类的社交网络,我们证明了现代有用的稀疏PCA在探索多元数据。R包epca和本文的补充材料可以在线获得。
Abstract Previous versions of sparse principal component analysis (PCA) have presumed that the eigen-basis (a p × k matrix) is approximately sparse. We propose a method that presumes the p × k matrix becomes approximately sparse after a k × k rotation. The simplest version of the algorithm initializes with the leading k principal components. Then, the principal components are rotated with an k × k orthogonal rotation to make them approximately sparse. Finally, soft-thresholding is applied to the rotated principal components. This approach differs from prior approaches because it uses an orthogonal rotation to approximate a sparse basis. One consequence is that a sparse component need not to be a leading eigenvector, but rather a mixture of them. In this way, we propose a new (rotated) basis for sparse PCA. In addition, our approach avoids “deflation” and multiple tuning parameters required for that. Our sparse PCA framework is versatile; for example, it extends naturally to a two-way analysis of a data matrix for simultaneous dimensionality reduction of rows and columns. We provide evidence showing that for the same level of sparsity, the proposed sparse PCA method is more stable and can explain more variance compared to alternative methods. Through three applications—sparse coding of images, analysis of transcriptome sequencing data, and large-scale clustering of social networks, we demonstrate the modern usefulness of sparse PCA in exploring multivariate data. An R package, epca, and the supplementary materials for this article are available online.