H-PARAFAC: Hierarchical Parallel Factor Analysis of Multidimensional Big Data

H-PARAFAC: Hierarchical Parallel Factor Analysis of Multidimensional Big Data
复制标题

H-PARAFAC:多维大数据的分层并行因子分析

DOI:
10.1109/tpds.2016.2613054
复制
发表时间:
2017-04-01
影响因子:
5.3
通讯作者:
Li, Xiaoli
Li, Xiaoli
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chen, Dan;Hu, Yangyang;Li, Xiaoli

文献摘要

被引文献

相似文献

如何从大量的多维数据中提取出隐含的多向因子,对叠加了大量噪声和干扰的数据进行分析,一直是各学科研究的重要课题。随着大数据时代数据规模和维度的快速增加,研究挑战出现,以(1)反映大张量的动态,同时在因子分解过程中不引入显著的失真,以及(2)处理复杂应用中噪声的影响。基于并行因子分析(PARAFAC)的“分治”理论,提出了一种基于GPU集群的分层并行处理框架H-PARAFAC. H-PARAFAC框架结合了用于协调子张量处理的粗粒度模型和用于计算每个子张量和融合子因子的细粒度并行模型。实验结果表明:(1)该方法突破了待分解多维数据规模的限制,在可扩展性和效率方面都明显优于传统方法,当数据量以<inline-formula><tex-math notation="LaTeX">n^3</tex-math><alternatives><inline-graphic xlink:href="wang-ieq1-2613054.gif"/></alternatives></inline-formula><inline-formula><tex-math notation="LaTeX">$的</tex-math><alternatives><inline-graphic xlink:href="wang-ieq2-2613054.gif"/></alternatives></inline-formula>数量级增加时,运行时间以n^2 $的数量级增加,(2)H-PARAFAC在抑制显著噪声的影响方面具有潜力,(3)H-PARAFAC在保持大张量的多模式特征方面远远上级传统的基于窗口的对应物。
It has long been an important issue in various disciplines to examine massive multidimensional data superimposed by a high level of noises and interferences by extracting the embedded multi-way factors. With the quick increases of data scales and dimensions in the big data era, research challenges arise in order to (1) reflect the dynamics of large tensors while introducing no significant distortions in the factorization procedure and (2) handle influences of the noises in sophisticated applications. A hierarchical parallel processing framework over a GPU cluster, namely H-PARAFAC, has been developed to enable scalable factorization of large tensors upon a “divide-and-conquer” theory for Parallel Factor Analysis (PARAFAC). The H-PARAFAC framework incorporates a coarse-grained model for coordinating the processing of sub-tensors and a fine-grained parallel model for computing each sub-tensor and fusing sub-factors. Experimental results indicate that (1) the proposed method breaks the limitation on the scale of multidimensional data to be factorized and dramatically outperforms the traditional counterparts in terms of both scalability and efficiency, e.g., the runtime increases in the order of <inline-formula> <tex-math notation="LaTeX">$n^2$</tex-math><alternatives><inline-graphic xlink:href="wang-ieq1-2613054.gif"/> </alternatives></inline-formula> when the data volume increases in the order of <inline-formula> <tex-math notation="LaTeX">$n^3$</tex-math><alternatives><inline-graphic xlink:href="wang-ieq2-2613054.gif"/> </alternatives></inline-formula>, (2) H-PARAFAC has potentials in refraining the influences of significant noises, and (3) H-PARAFAC is far superior to the conventional window-based counterparts in preserving the features of multiple modes of large tensors.