An implementation of a randomized algorithm for principal component analysis

An implementation of a randomized algorithm for principal component analysis
复制标题

DOI:
--
复制
发表时间:
2014-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Arthur Szlam;Y. Kluger;M. Tygert
Arthur Szlam;Y. Kluger;M. Tygert
中科院分区:
其他
文献类型:
--
作者:
Arthur Szlam;Y. Kluger;M. Tygert

文献摘要

被引文献

相似文献

近年来,低秩近似的随机化方法得到了迅速的发展。这些方法的目标是主成分分析(PCA)和截断奇异值分解(SVD)的计算。本文提出了一个基本上黑盒,傻瓜的实现Mathworks的MATLAB,一个流行的软件平台,数值计算。如通过几个测试所示,用于低秩近似的随机化算法在基本上所有方面都优于或至少匹配经典技术(例如Lanczos迭代):准确性,计算效率(速度和内存使用),易用性,并行性和可靠性。然而,经典方法仍然是估计谱范数的首选方法,并且在计算最小奇异值和相应的奇异向量(或奇异子空间)方面要上级得多。
Recent years have witnessed intense development of randomized methods for low-rank approximation. These methods target principal component analysis (PCA) and the calculation of truncated singular value decompositions (SVD). The present paper presents an essentially black-box, fool-proof implementation for Mathworks' MATLAB, a popular software platform for numerical computation. As illustrated via several tests, the randomized algorithms for low-rank approximation outperform or at least match the classical techniques (such as Lanczos iterations) in basically all respects: accuracy, computational efficiency (both speed and memory usage), ease-of-use, parallelizability, and reliability. However, the classical procedures remain the methods of choice for estimating spectral norms, and are far superior for calculating the least singular values and corresponding singular vectors (or singular subspaces).