Designing, Implementing, and Evaluating the Upcoming OpenSHMEM Teams API

Designing, Implementing, and Evaluating the Upcoming OpenSHMEM Teams API
复制标题

设计、实现和评估即将推出的 OpenSHMEM Teams API

DOI:
--
复制
发表时间:
2019
期刊:
2019 IEEE/ACM Parallel Applications Workshop, Alternatives To MPI (PAW-ATM)
影响因子:
--
通讯作者:
James Dinan
James Dinan
中科院分区:
--
文献类型:
--
作者:
David Ozog;Md. Wasi;Gerard Taylor;James Dinan

文献摘要

被引文献

相似文献

多年来,OpenSHMEM并行编程接口提供了MPI的高性能替代方案,它强调单侧消息传递,简化了跨全局内存空间的通信,并增强了快速发展的结构互连技术的能力。OpenSHMEM规范标准化了库接口,优先考虑高性能和可移植的API。在几个权威供应商和研究人员的大力支持下,该规范不断成熟。例如,OpenSHMEM规范委员会正在积极标准化一个独特的团队API,该API使用户定义的应用程序进程子集能够高效地执行通信操作,如集合例程、远程内存访问和远程原子操作。本文描述了OpenSHMEM团队接口、实现API的几个有趣的方面和挑战,以及可以提高可编程性和/或性能的可能扩展。我们评估了初步实现的性能,并表明有效地使用团队可以促进大规模集体操作的令人印象深刻的改进(2-16倍的加速,取决于算法,节点计数和缓冲区大小),即使在简化底层编程模型的同时。
For many years, the OpenSHMEM parallel pro- gramming interface has provided a high-performance alternative to MPI that emphasizes one-sided messaging, simplifies com- munication across a global memory space, and bolsters the capabilities of rapidly evolving fabric interconnect technolo- gies. The OpenSHMEM specification standardizes the library interfaces, prioritizing a performant and portable API. The specification continues to mature with vigorous support from several authoritative vendors and researchers. For example, the OpenSHMEM specification committee is actively standardizing a unique teams API that enables user-defined subsets of application processes to efficiently and productively perform communication operations, such as collectives routines, remote memory accesses, and remote atomic operations. This paper describes the OpenSHMEM teams interface and several interesting aspects and challenges in implementing the API, as well as possible extensions that could improve the pro- grammability and/or performance. We evaluate the performance of a preliminary implementation and show that using teams effectively can facilitate impressive improvements of collective operations at scale (2-16x speedup, depending on the algorithm, node count, and buffer size), even while simplifying the under- lying programming model.