An intuitive and most efficient Ll-norm principal component analysis algorithm for big data

An intuitive and most efficient Ll-norm principal component analysis algorithm for big data
复制标题

一种直观且最高效的大数据Ll范数主成分分析算法

DOI:
--
复制
发表时间:
2019
期刊:
Annual Conference on Information Sciences and Systems
影响因子:
--
通讯作者:
Xiaowei Song
Xiaowei Song
中科院分区:
--
文献类型:
--
作者:
Xiaowei Song

文献摘要

被引文献

相似文献

格拉斯曼平均(GA)可以与l范数主成分(PC)重合,并且在数百万个样本中具有可扩展性。然而,目前尚不清楚是否存在,以及通过修改基于定点优化的遗传算法可以获得多少进一步的速度提高。在本文中,我对这一优化过程进行了直观的分析,并提出了改进方案,即一种不需要任何迭代的在线算法。我表明,它可以是最有效的,因为它每台PC只访问每个样本一次,内存需求最小,不像GA或MATLAB svds。事实证明,对于大数据,它是收敛的。
Grassmann average (GA) can coincide with Ll- norm principal component (PC) and is scalable for millions of samples. However, it is unclear whether there exists and how much further speed improvement can be gained by revising the fixed-point optimization-based GA. In this paper, I analyze such optimization process in an intuitive way and propose its improvement, i.e., an online algorithm without any iterations. I show that it can be most efficient in the sense that it only visits each sample once per PC, with minimal memory requirement, unlike GA or MATLAB svds. It is proved to be convergent for big data.