Improving the Scaling of an Asynchronous Many-Task Runtime with a Lightweight Communication Engine

Improving the Scaling of an Asynchronous Many-Task Runtime with a Lightweight Communication Engine
复制标题

使用轻量级通信引擎改进异步多任务运行时的扩展

DOI:
10.1145/3605573.3605642
复制
发表时间:
2023
期刊:
ACM
影响因子:
--
通讯作者:
Snir, Marc
Snir, Marc
中科院分区:
--
文献类型:
--
作者:
Mor, Omri;Bosilca, George;Snir, Marc

文献摘要

参考文献

被引文献

相似文献

异步多任务(AMT)运行时作为一种有效的方式映射到异构计算资源的不规则和动态的并行应用程序的兴趣越来越大。在这项工作中,我们表明,AMT仍然与通信瓶颈的斗争时,强大的扩展计算和常用的通信库,如MPI的设计有助于这些瓶颈。我们用LCI代替MPI,LCI是一种轻量级通信接口,专为动态异步框架设计,作为PaRSEC运行时的通信层。其结果是通信微基准测试中的端到端延迟显著减少,并且在HiCMA(一种基于图块的低秩Cholesky分解包)中将整体求解时间减少了12%。
There is a growing interest in Asynchronous Many-Task (AMT) runtimes as an efficient way to map irregular and dynamic parallel applications onto heterogeneous computing resources. In this work, we show that AMTs nonetheless struggle with communication bottlenecks when scaling computations strongly and that the design of commonly-used communication libraries such as MPI contribute to these bottlenecks. We replace MPI with LCI, a Lightweight Communication Interface that is designed for dynamic, asynchronous frameworks, as the communication layer for the PaRSEC runtime. The result is a significant reduction of end-to-end latency in communication microbenchmarks and a reduction of overall time-to-solution by up to 12% in HiCMA, a tile-based low-rank Cholesky factorization package.
DOI: 10.1109/ipdpsw52791.2021.00079
发表时间: 2021-06
期刊: 2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)
影响因子: --
作者:
Jaemin Choi;Zane Fink;Sam White;Nitin Bhat;D. Richards;L. Kalé
通讯作者: Jaemin Choi;Zane Fink;Sam White;Nitin Bhat;D. Richards;L. Kalé
DOI: 10.1109/pact.2019.00010
发表时间: 2019-09
期刊: 2019 28th International Conference on Parallel Architectures and Compilation Techniques (PACT)
影响因子: --
作者:
Roshan Dathathri;G. Gill;Loc Hoang;Vishwesh Jatala;K. Pingali;V. K. Nandivada;Hoang-Vu Dang;M. Snir-M.
通讯作者: Roshan Dathathri;G. Gill;Loc Hoang;Vishwesh Jatala;K. Pingali;V. K. Nandivada;Hoang-Vu Dang;M. Snir-M.
使用 DPLASMA 在大规模并行架构上灵活开发密集线性代数算法
DOI: --
发表时间: 2011
期刊: IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子: --
作者:
G. Bosilca;A. Bouteiller;Anthony Danalis;Mathieu Faverge;A. Haidar;T. Hérault;J. Kurzak;J. Langou;Pierre Lemarinier;H. Ltaief;P. Luszczek;A. YarKhan;J. Dongarra
通讯作者: J. Dongarra
MPI 中的分层时钟同步
DOI: --
发表时间: 2018
期刊: IEEE International Conference on Cluster Computing
影响因子: --
作者:
S. Hunold;Alexandra Carpen
通讯作者: Alexandra Carpen
给 MPI 线程一个公平的机会:多线程 MPI 设计研究
DOI: 10.1109/cluster.2019.8891015
发表时间: 2019
期刊: IEEE Cluster
影响因子: --
作者:
Patinyasakdikul, T.;Eberius, D.;Bosilca, G.;Hjelm, N.
通讯作者: Hjelm, N.