fgSpMSpV: A Fine-grained Parallel SpMSpV Framework on HPC Platforms

fgSpMSpV: A Fine-grained Parallel SpMSpV Framework on HPC Platforms
复制标题

fgSpMSpV:HPC 平台上的细粒度并行 SpMSpV 框架

DOI:
10.1145/3512770
复制
发表时间:
2022-04
期刊:
ACM Transactions on Parallel Computing (ACM并行计算领域旗舰期刊)
影响因子:
--
通讯作者:
Albert Y. Zomaya
Albert Y. Zomaya
中科院分区:
其他
文献类型:
--
作者:
Yuedan Chen;Guoqing Xiao;Kenli Li;Francesco Piccialli;Albert Y. Zomaya

文献摘要

参考文献

被引文献

相似文献

稀疏矩阵-稀疏向量(Sparse matrix-sparse vector,SpMSpV)乘法是许多高性能科学和工程应用中的基本和重要运算之一。固有的不规则性和较差的数据局部性导致在高性能计算(HPC)系统上扩展SpMSpV的两个主要挑战:(i)大量冗余数据限制了带宽和并行资源的利用;(ii)不规则的访问模式限制了计算资源的利用。本文提出了一个细粒度的并行SpMSpV(fgSpMSpV)框架上的神威太湖之光超级计算机,以减轻大规模的现实世界中的应用的挑战。首先,fgSpMSpV采用MPI \(+ \)OpenMP \(+X \)并行化模型,充分利用异构HPC体系结构的多级混合并行性,加速前/后处理和主要的SpMSpV计算。其次,fgSpMSpV利用自适应并行执行来减少预处理,适应Sunway系统的并行性和内存层次结构,同时仍然驯服SpMSpV中的冗余和随机内存访问,包括一组技术,如细粒度分区器,重新收集方法和压缩稀疏列向量(CSCV)矩阵格式。第三,fgSpMSpV使用几种优化技术来进一步利用计算资源。在不同的输入稀疏度下,神威太湖之光上的fgSpMSpV通过关键的优化技术获得了显著的性能提升。此外,fgSpMSpV在NVIDIA Tesal P100 GPU上实现,并应用于呼吸优先搜索(BFS)应用程序。在P100 GPU上的fgSpMSpV获得了最高达134.38倍的最新SpMSpV算法的加速比,使用fgSpMSpV的BFS应用程序获得了最高达21.68倍的最新SpMSpV算法的加速比。
Sparse matrix-sparse vector (SpMSpV) multiplication is one of the fundamental and important operations in many high-performance scientific and engineering applications. The inherent irregularity and poor data locality lead to two main challenges to scaling SpMSpV over high-performance computing (HPC) systems: (i) a large amount of redundant data limits the utilization of bandwidth and parallel resources; (ii) the irregular access pattern limits the exploitation of computing resources. This paper proposes a fine-grained parallel SpMSpV (fgSpMSpV) framework on Sunway TaihuLight supercomputer to alleviate the challenges for large-scale real-world applications. First, fgSpMSpV adopts an MPI \( + \) OpenMP \( +X \) parallelization model to exploit the multi-stage and hybrid parallelism of heterogeneous HPC architectures and accelerate both pre-/post-processing and main SpMSpV computation. Second, fgSpMSpV utilizes an adaptive parallel execution to reduce the pre-processing, adapt to the parallelism and memory hierarchy of the Sunway system, while still tame redundant and random memory accesses in SpMSpV, including a set of techniques like the fine-grained partitioner, re-collection method, and Compressed Sparse Column Vector (CSCV) matrix format. Third, fgSpMSpV uses several optimization techniques to further utilize the computing resources. fgSpMSpV on the Sunway TaihuLight gains a noticeable performance improvement from the key optimization techniques with various sparsity of the input. Additionally, fgSpMSpV is implemented on an NVIDIA Tesal P100 GPU and applied to the breath-first-search (BFS) application. fgSpMSpV on a P100 GPU obtains the speedup of up to \( 134.38\times \) over the state-of-the-art SpMSpV algorithms, and the BFS application using fgSpMSpV achieves the speedup of up to \( 21.68\times \) over the state-of-the-arts.
DOI: 10.1109/tpds.2018.2848618
发表时间: 2018-06
影响因子: 5.3
作者:
Lixin He;Hong An;Chao Yang;Fei Wang;Junshi Chen;Chao Wang;Weihao Liang;Shaojun Dong;Qiao S
通讯作者: Lixin He;Hong An;Chao Yang;Fei Wang;Junshi Chen;Chao Wang;Weihao Liang;Shaojun Dong;Qiao S
DOI: 10.1109/ipdps.2013.52
发表时间: 2013-05
期刊: 2013 IEEE 27th International Symposium on Parallel and Distributed Processing
影响因子: --
作者:
A. Buluç;Erika Duriakova;A. Fox;J. Gilbert;Shoaib Kamil;A. Lugowski;L. Oliker;Samuel Williams
通讯作者: A. Buluç;Erika Duriakova;A. Fox;J. Gilbert;Shoaib Kamil;A. Lugowski;L. Oliker;Samuel Williams
DOI: 10.23919/date.2019.8714836
发表时间: 2019-03
期刊: 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子: --
作者:
Alwin Zulehner;R. Wille
通讯作者: Alwin Zulehner;R. Wille
DOI: 10.1109/tpds.2019.2906166
发表时间: 2019-10
影响因子: 5.3
作者:
Muhammet Mustafa Ozdal
通讯作者: Muhammet Mustafa Ozdal
DOI: 10.14778/2809974.2809983
发表时间: 2015-03
期刊: ArXiv
影响因子: --
作者:
N. Sundaram;N. Satish;Md. Mostofa Ali Patwary;Subramanya R. Dulloor;Michael J. Anderson;Satya Gautam Vadlamudi-Satya-Gautam-V
通讯作者: N. Sundaram;N. Satish;Md. Mostofa Ali Patwary;Subramanya R. Dulloor;Michael J. Anderson;Satya Gautam Vadlamudi-Satya-Gautam-V