Alternating maximization: unifying framework for 8 sparse PCA formulations and efficient parallel codes

Alternating maximization: unifying framework for 8 sparse PCA formulations and efficient parallel codes
复制标题

交替最大化:8 个稀疏 PCA 公式和高效并行代码的统一框架

DOI:
10.1007/s11081-020-09562-3
复制
发表时间:
2020
影响因子:
2.1
通讯作者:
Takáč, Martin
Takáč, Martin
中科院分区:
工程技术3区
文献类型:
--
作者:
Richtárik, Peter;Jahani, Majid;Ahipaşaoğlu, Selin Damla;Takáč, Martin

文献摘要

参考文献

被引文献

相似文献

给定一个多变量数据集,稀疏主成分分析(SPCA)的目的是提取几个变量的线性组合,这些变量尽可能多地解释数据中的方差,同时控制这些组合中非零载荷的数量。在本文中,我们考虑8种不同的优化公式计算一个单一的稀疏加载向量:我们采用两个标准的测量方差(L2,L1)和两个稀疏诱导规范(L0,L1),这是在两种方式(约束,惩罚)。我们的配方,特别是一个与L0约束和L1方差,没有被认为是在文献中。我们给出了一个统一的重新制定,我们建议通过交替最大化(AM)方法来解决。我们表明,AM是等价于GPower的所有配方。除此之外,我们还提供了24个高效的并行SPCA实现:8个问题中的每个问题有3个代码(多核,GPU和集群)。方法中的并行性旨在(1)加速计算(我们的GPU代码可以比用C++编写的高效串行代码快100倍),(2)获得解释更多方差的解决方案,(3)处理大数据问题(我们的集群代码可以在一分钟内解决357 GB的问题)。
Given a multivariate data set, sparse principal component analysis (SPCA) aims to extract several linear combinations of the variables that together explain the variance in the data as much as possible, while controlling the number of nonzero loadings in these combinations. In this paper we consider 8 different optimization formulations for computing a single sparse loading vector: we employ two norms for measuring variance (L2, L1) and two sparsity-inducing norms (L0, L1), which are used in two ways (constraint, penalty). Three of our formulations, notably the one with L0 constraint and L1 variance, have not been considered in the literature. We give a unifying reformulation which we propose to solve via the alternating maximization (AM) method. We show that AM is equivalent to GPower for all formulations. Besides this, we provide 24 efficient parallel SPCA implementations: 3 codes (multi-core, GPU and cluster) for each of the 8 problems. Parallelism in the methods is aimed at (1) speeding up computations (our GPU code can be 100 times faster than an efficient serial code written in C++), (2) obtaining solutions explaining more variance and (3) dealing with big data problems (our cluster code can solve a 357 GB problem in a minute).
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者:
T. Bouwmans;N. Aybat;E. Zahzah
通讯作者: E. Zahzah
寻找极端特征向量的稀疏近似:稀疏 PCA 和扩展的广义幂方法
DOI: --
发表时间: 2011
期刊:
影响因子: --
作者:
Peter Richtárik
通讯作者: Peter Richtárik
DOI: 10.1093/biostatistics/kxp008
发表时间: 2009-07-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Witten, Daniela M.;Tibshirani, Robert;Hastie, Trevor
通讯作者: Hastie, Trevor