FAST-PCA: A Fast and Exact Algorithm for Distributed Principal Component Analysis

FAST-PCA: A Fast and Exact Algorithm for Distributed Principal Component Analysis
复制标题

快速主成分分析(FAST - PCA):一种用于分布式主成分分析的快速精确算法

DOI:
10.1109/tsp.2022.3229635
复制
发表时间:
2021-08
影响因子:
5.4
通讯作者:
Arpita Gang;W. Bajwa
Arpita Gang;W. Bajwa
中科院分区:
工程技术1区
文献类型:
--
作者:
Arpita Gang;W. Bajwa

文献摘要

相似文献

主成分分析(PCA)是机器学习领域的基本数据预处理工具。虽然PCA通常被认为是一种降维方法,但PCA的目的实际上是双重的:降维和不相关特征学习。此外,现代数据集的维度和样本量的巨大使得集中的PCA解决方案无法使用。在这种情况下,本文重新考虑了当数据样本分布在任意连接的网络中跨节点时的主成分分析问题。虽然存在一些分布式PCA的解决方案,但这些解决方案要么忽略了PCA的不相关特征学习方面,往往具有较高的通信开销,使它们效率低下,或者缺乏“精确”或“全局”收敛保证。为了克服上述问题,本文提出了一种名为Fast -PCA (Fast and exAct distributed PCA)的分布式PCA算法。该算法在通信方面是有效的,并且被证明是线性和精确收敛于主成分,导致降维和不相关的特征。实验结果进一步支持了这些说法。
Principal Component Analysis (PCA) is a fundamental data preprocessing tool in the world of machine learning. While PCA is often thought of as a dimensionality reduction method, the purpose of PCA is actually two-fold: dimension reduction and uncorrelated feature learning. Furthermore, the enormity of the dimensions and sample size in the modern day datasets have rendered the centralized PCA solutions unusable. In that vein, this paper reconsiders the problem of PCA when data samples are distributed across nodes in an arbitrarily connected network. While a few solutions for distributed PCA exist, those either overlook the uncorrelated feature learning aspect of the PCA, tend to have high communication overhead that makes them inefficient and/or lack ‘exact’ or ‘global’ convergence guarantees. To overcome these aforementioned issues, this paper proposes a distributed PCA algorithm termed FAST-PCA (Fast and exAct diSTributed PCA). The proposed algorithm is efficient in terms of communication and is proven to converge linearly and exactly to the principal components, leading to dimension reduction as well as uncorrelated features. The claims are further supported by experimental results.