Generalized Flow-Graph Programming Using Template Task-Graphs: Initial Implementation and Assessment

Generalized Flow-Graph Programming Using Template Task-Graphs: Initial Implementation and Assessment
复制标题

DOI:
10.1109/ipdps53621.2022.00086
复制
发表时间:
2022-05
期刊:
2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
J. Schuchart;Poornima Nookala;M. Javanmard;T. Hérault;Edward F. Valeev;G. Bosilca;R. Harrison
J. Schuchart;Poornima Nookala;M. Javanmard;T. Hérault;Edward F. Valeev;G. Bosilca;R. Harrison
中科院分区:
其他
文献类型:
--
作者:
J. Schuchart;Poornima Nookala;M. Javanmard;T. Hérault;Edward F. Valeev;G. Bosilca;R. Harrison

文献摘要

相似文献

我们提出并评价了TTG,一种新颖的编程模型及其c++实现,它结合了控制和数据流程图编程的思想,支持紧凑的规范和高效的动态和不规则应用程序的分布式执行。支持基于任务执行的编程接口通常只支持共享内存并行环境;一些支持分布式内存环境,或者通过发现所有进程上的任务的整个DAG,或者通过引入显式通信。第一种方法限制了可伸缩性,而第二种方法增加了编程的复杂性。我们将演示TTG如何在不牺牲可伸缩性或可编程性的情况下解决这些问题,方法是提供比以任务为中心的编程系统通常提供的更高级别的抽象,而不会妨碍这些运行时有效地管理任务创建和执行以及数据和资源管理的能力。TTG支持在2个不同的任务运行时(PaRSEC和MADNESS)上执行分布式内存。在TTG中实现的具有不同程度不规则性的四种典型应用程序(图分析、密集和块稀疏线性代数以及数值积分微分)的性能在大型分布式内存平台上进行了说明,并与最先进的实现进行了比较。
We present and evaluate TTG, a novel programming model and its C++ implementation that by marrying the ideas of control and data flowgraph programming supports compact specification and efficient distributed execution of dynamic and irregular applications. Programming interfaces that support task-based execution often only support shared memory parallel environments; a few support distributed memory environments, either by discovering the entire DAG of tasks on all processes, or by introducing explicit communications. The first approach limits scalability, while the second increases the complexity of programming. We demonstrate how TTG can address these issues without sacrificing scalability or programmability by providing higher-level abstractions than conventionally provided by task-centric programming systems, without impeding the ability of these runtimes to manage task creation and execution as well as data and resource management efficiently. TTG supports distributed memory execution over 2 different task runtimes, PaRSEC and MADNESS. Performance of four paradigmatic applications (in graph analytics, dense and block-sparse linear algebra, and numerical integrodifferential calculus) with various degrees of irregularity implemented in TTG is illustrated on large distributed-memory platforms and compared to the state-of-the-art implementations.