The Best of Both Worlds: Distributed PCA That is Both Exact and Communication Efficient

The Best of Both Worlds: Distributed PCA That is Both Exact and Communication Efficient
复制标题

DOI:
10.23919/eusipco55093.2022.9909543
复制
发表时间:
2022-08
期刊:
2022 30th European Signal Processing Conference (EUSIPCO)
影响因子:
--
通讯作者:
Arpita Gang;W. Bajwa
Arpita Gang;W. Bajwa
中科院分区:
其他
文献类型:
--
作者:
Arpita Gang;W. Bajwa

文献摘要

相似文献

机器学习算法的有效性在很大程度上取决于数据表示的良好性。虽然现代数据在维度和数量上的连续性需要降维和特征提取以有效使用可用的计算资源,但是已知使用不相关的特征来增强这种机器学习算法的性能。因此,一个有效的表示学习方法应该集中在降维以及不相关的特征提取。尽管主成分分析(PCA)和线性自编码器是主要用于降维的基本数据处理工具,但如果设计得当,它们也可以用于提取不相关的特征。与此同时,不断增加的数据量或固有的分布式数据生成等因素阻碍了现有集中式解决方案的使用,这些解决方案需要在单个位置提供数据。本文提出了一种算法的两个变种,称为FAST-PCA(快速和exAct diSTributed PCA)基于前馈神经网络的系统,在分布式设置中学习数据表示,使它们在维度上减少,以及具有不相关的功能。所提出的变体是为了抑制现有解决方案中普遍存在的通信开销,并以线性速率收敛到精确的解决方案。这些主张得到了广泛的数值实验的进一步支持。
The effectiveness of machine learning algorithms largely depends on the goodness of the representation of data. While the massiveness in dimension and amount of modern day data requires dimension reduction and feature extraction for efficient use of available computational resources, the use of un-correlated features is known to enhance the performance of such machine learning algorithms. Thus, an efficient representation learning approach should focus on dimension reduction as well as uncorrelated feature extraction. Even though Principal Component Analysis (PCA) and linear autoencoders are fundamental data processing tools largely used for dimension reduction, they can also be used to extract uncorrelated features when engineered properly. At the same time, factors like ever-increasing volume of data or inherently distributed data generation impede the use of existing centralized solutions for representation learning that require availability of data at a single location. This paper proposes two variants of an algorithm called FAST-PCA (Fast and exAct diSTributed PCA) based on a feedforward neural network-based system that learn data representations in a distributed setting such that they are reduced in dimension as well as have uncorrelated features. The proposed variants are meant to curb the communication overheads prevalent in the existing solutions and are shown to converge to the exact solutions at a linear rate. These claims are further supported by extensive numerical experiments.