fgSpMSpV: A Fine-grained Parallel SpMSpV Framework on HPC Platforms
fgSpMSpV: A Fine-grained Parallel SpMSpV Framework on HPC Platforms
复制标题
fgSpMSpV:HPC 平台上的细粒度并行 SpMSpV 框架
DOI:
10.1145/3512770
复制
发表时间:
2022-04
期刊:
影响因子:
--
通讯作者:
Albert Y. Zomaya
中科院分区:
文献类型:
--
作者:
Yuedan Chen;Guoqing Xiao;Kenli Li;Francesco Piccialli;Albert Y. Zomaya
Sparse matrix-sparse vector (SpMSpV) multiplication is one of the fundamental and important operations in many high-performance scientific and engineering applications. The inherent irregularity and poor data locality lead to two main challenges to scaling SpMSpV over high-performance computing (HPC) systems: (i) a large amount of redundant data limits the utilization of bandwidth and parallel resources; (ii) the irregular access pattern limits the exploitation of computing resources. This paper proposes a fine-grained parallel SpMSpV (fgSpMSpV) framework on Sunway TaihuLight supercomputer to alleviate the challenges for large-scale real-world applications. First, fgSpMSpV adopts an MPI \( + \) OpenMP \( +X \) parallelization model to exploit the multi-stage and hybrid parallelism of heterogeneous HPC architectures and accelerate both pre-/post-processing and main SpMSpV computation. Second, fgSpMSpV utilizes an adaptive parallel execution to reduce the pre-processing, adapt to the parallelism and memory hierarchy of the Sunway system, while still tame redundant and random memory accesses in SpMSpV, including a set of techniques like the fine-grained partitioner, re-collection method, and Compressed Sparse Column Vector (CSCV) matrix format. Third, fgSpMSpV uses several optimization techniques to further utilize the computing resources. fgSpMSpV on the Sunway TaihuLight gains a noticeable performance improvement from the key optimization techniques with various sparsity of the input. Additionally, fgSpMSpV is implemented on an NVIDIA Tesal P100 GPU and applied to the breath-first-search (BFS) application. fgSpMSpV on a P100 GPU obtains the speedup of up to \( 134.38\times \) over the state-of-the-art SpMSpV algorithms, and the BFS application using fgSpMSpV achieves the speedup of up to \( 21.68\times \) over the state-of-the-arts.
登录
查看更多内容
DOI:
10.1109/tpds.2018.2848618
发表时间:
2018-06
影响因子:
5.3
作者:
Lixin He;Hong An;Chao Yang;Fei Wang;Junshi Chen;Chao Wang;Weihao Liang;Shaojun Dong;Qiao S
通讯作者:
Lixin He;Hong An;Chao Yang;Fei Wang;Junshi Chen;Chao Wang;Weihao Liang;Shaojun Dong;Qiao S
DOI:
10.1109/ipdps.2013.52
发表时间:
2013-05
期刊:
2013 IEEE 27th International Symposium on Parallel and Distributed Processing
影响因子:
--
作者:
A. Buluç;Erika Duriakova;A. Fox;J. Gilbert;Shoaib Kamil;A. Lugowski;L. Oliker;Samuel Williams
通讯作者:
A. Buluç;Erika Duriakova;A. Fox;J. Gilbert;Shoaib Kamil;A. Lugowski;L. Oliker;Samuel Williams
DOI:
10.23919/date.2019.8714836
发表时间:
2019-03
期刊:
2019 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
作者:
Alwin Zulehner;R. Wille
通讯作者:
Alwin Zulehner;R. Wille
DOI:
10.1109/tpds.2019.2906166
发表时间:
2019-10
影响因子:
5.3
作者:
Muhammet Mustafa Ozdal
通讯作者:
Muhammet Mustafa Ozdal
DOI:
10.14778/2809974.2809983
发表时间:
2015-03
期刊:
ArXiv
影响因子:
--
作者:
N. Sundaram;N. Satish;Md. Mostofa Ali Patwary;Subramanya R. Dulloor;Michael J. Anderson;Satya Gautam Vadlamudi-Satya-Gautam-V
通讯作者:
N. Sundaram;N. Satish;Md. Mostofa Ali Patwary;Subramanya R. Dulloor;Michael J. Anderson;Satya Gautam Vadlamudi-Satya-Gautam-V