DISTRIBUTED ESTIMATION OF PRINCIPAL EIGENSPACES

DISTRIBUTED ESTIMATION OF PRINCIPAL EIGENSPACES
复制标题

DOI:
10.1214/18-aos1713
复制
发表时间:
2019-12-01
影响因子:
4.5
通讯作者:
Zhu, Ziwei
Zhu, Ziwei
中科院分区:
数学1区
文献类型:
--
作者:
Fan, Jianqing;Wang, Dong;Zhu, Ziwei

文献摘要

被引文献

相似文献

主成分分析(PCA)是统计机器学习的基础。它提取了导致数据变化最大的潜在主要因素。然而,当数据存储在多台机器上时,由于通信成本的限制,PCA无法集中计算,因此需要分布式PCA算法。本文提出并研究了一种分布式PCA算法:每个节点机计算最上面的K个特征向量,并将其传输到中央服务器;然后,中央服务器聚合来自所有节点机器的信息,并基于聚合的信息执行PCA。我们研究了得到的最上面K个特征向量的分布估计量的偏差和方差。特别是,我们证明了对于具有对称创新的分布,经验顶特征空间是无偏的,因此分布式PCA是“无偏的”。我们推导了分布式PCA估计的收敛率,它明确地依赖于协方差的有效秩、特征集和机器的数量。我们表明,当机器数量不是不合理的大时,即使没有对整个数据的完全访问,分布式主成分分析的性能也与整个样本主成分分析一样好。通过大量的仿真研究验证了理论结果。我们还将分析扩展到异质情况,其中总体协方差矩阵在局部机器上不同,但共享相似的顶部特征结构。
Principal component analysis (PCA) is fundamental to statistical machine learning. It extracts latent principal factors that contribute to the most variation of the data. When data are stored across multiple machines, however, communication cost can prohibit the computation of PCA in a central location and distributed algorithms for PCA are thus needed. This paper proposes and studies a distributed PCA algorithm: each node machine computes the top K eigenvectors and transmits them to the central server; the central server then aggregates the information from all the node machines and conducts a PCA based on the aggregated information. We investigate the bias and variance for the resulting distributed estimator of the top K eigenvectors. In particular, we show that for distributions with symmetric innovation, the empirical top eigenspaces are unbiased, and hence the distributed PCA is "unbiased." We derive the rate of convergence for distributed PCA estimators, which depends explicitly on the effective rank of covariance, eigengap, and the number of machines. We show that when the number of machines is not unreasonably large, the distributed PCA performs as well as the whole sample PCA, even without full access of whole data. The theoretical results are verified by an extensive simulation study. We also extend our analysis to the heterogeneous case where the population covariance matrices are different across local machines but share similar top eigenstructures.