Sparse Tensor Factorization on Many-Core Processors with High-Bandwidth Memory

Sparse Tensor Factorization on Many-Core Processors with High-Bandwidth Memory
复制标题

具有高带宽内存的多核处理器上的稀疏张量分解

DOI:
10.1109/ipdps.2017.84
复制
发表时间:
2017
期刊:
2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
G. Karypis
G. Karypis
中科院分区:
--
文献类型:
--
作者:
Shaden Smith;Jongsoo Park;G. Karypis

文献摘要

被引文献

相似文献

HPC系统越来越多地用于数据密集型计算,其表现出不规则的存储器访问、非均匀的工作分布、大的存储器占用空间和高的存储器带宽需求。为了满足这些具有挑战性的需求,HPC系统正在转向众核架构,这些架构具有大量高能效内核,并由高带宽内存支持。这些功能在英特尔最近的Knights Landing众核处理器(KNL)中得到了体现,该处理器通常具有68个内核和16 GB的封装多通道DRAM(MCDRAM)。这项工作研究了如何由KNL提供的新的架构功能,可以使用在分解稀疏,非结构化张量的情况下,使用典型的polyadic分解(CPD)。CPD被广泛用于分析各个领域的大型多路数据集,包括精准医疗、网络安全和电子商务。为此,我们(一)开发的CPD是服从数百个并发线程的问题分解,同时保持负载平衡和低同步成本;及(ii)探索利用的架构功能,如MCDRAM。使用一个KNL处理器,我们的算法实现了高达1.8倍的加速比双插槽英特尔至强系统与44个核心。
HPC systems are increasingly used for data intensive computations which exhibit irregular memory accesses, non-uniform work distributions, large memory footprints, and high memory bandwidth demands. To address these challenging demands, HPC systems are turning to many-core architectures that feature a large number of energy-efficient cores backed by high-bandwidth memory. These features are exemplified in Intel's recent Knights Landing many-core processor (KNL), which typically has 68 cores and 16GB of on-package multi-channel DRAM (MCDRAM). This work investigates how the novel architectural features offered by KNL can be used in the context of decomposing sparse, unstructured tensors using the canonical polyadic decomposition (CPD). The CPD is used extensively to analyze large multi-way datasets arising in various areas including precision healthcare, cybersecurity, and e-commerce. Towards this end, we (i) develop problem decompositions for the CPD which are amenable to hundreds of concurrent threads while maintaining load balance and low synchronization costs; and (ii) explore the utilization of architectural features such as MCDRAM. Using one KNL processor, our algorithm achieves up to 1.8x speedup over a dual socket Intel Xeon system with 44 cores.