Low-cost MPI Multithreaded Message Matching Benchmarking

Low-cost MPI Multithreaded Message Matching Benchmarking
复制标题

低成本 MPI 多线程消息匹配基准测试

DOI:
10.1109/hpcc-smartcity-dss50907.2020.00022
复制
发表时间:
2020
期刊:
2020 IEEE 22nd International Conference on High Performance Computing and Communications; IEEE 18th International Conference on Smart City; IEEE 6th International Conference on Data Science and Systems (HPCC/SmartCity/DSS)
影响因子:
--
通讯作者:
William Marts
William Marts
中科院分区:
--
文献类型:
--
作者:
William Schonbein;Ryan E. Grant;Scott Levy;Matthew G. F. Dosanjh;William Marts

文献摘要

参考文献

被引文献

相似文献

消息传递接口(MPI)标准允许用户级线程并发调用MPI库。虽然此功能目前很少使用,但开发人员对在不久的将来采用它有相当大的兴趣。有理由相信,多线程通信可能会导致额外的消息处理开销,在解复用期间搜索的项目数量和搜索所花费的时间量,因为它有可能增加交换的消息数量,并引入不确定的消息排序。因此,理解向MPI应用程序添加多线程的含义对于未来的应用程序开发非常重要。推进这种理解的一个策略是通过“低成本”基准测试,使用更少的资源来模拟完整的通信模式。例如,虽然完整的“真实世界”多线程光环交换需要9个或27个节点,但低成本替代方案仅需要两个,使得其可部署在由于高利用率而难以获取资源的系统上(例如,忙碌的容量计算系统),或者因为不存在必要的资源而不可能(例如,节点太少的测试床)。虽然已经提出了这样的基准,但报告的结果仅限于单一架构或通过模拟间接得出,并且没有尝试确认低成本基准准确地捕获完整(非仿真)交换的特征。此外,基准代码还没有公开,本文提出的研究的目的是量化如何准确的低成本基准捕捉匹配行为的完整,真实世界的基准。在此过程中,我们还倡导低成本基准的可行性和实用性。我们提出了一个“现实世界”的基准实现一个完整的多线程光环交换9和27节点,定义为5点和9点的二维stenchmark,和7点和27点的三维stenchmark。同样,我们提出了一个“低成本”基准,仅使用两个节点来模拟这些通信模式。然后,我们确认,在多个架构,低成本的基准提供了准确的估计,在消息处理过程中搜索的项目数量,以及处理这些消息所花费的时间。最后,我们展示了低成本基准测试的实用性,通过使用它来分析最先进的Mellanox ConnectX-5硬件支持卸载MPI消息解复用的性能影响。为了便于进一步研究多线程MPI对消息匹配行为的影响,我们的两个基准测试的源代码将包含在Sandia MPI Micro-Benchmark Suite的下一个发布版本中。
The Message Passing Interface (MPI) standard allows user-level threads to concurrently call into an MPI library. While this feature is currently rarely used, there is considerable interest from developers in adopting it in the near future. There is reason to believe that multithreaded communication may incur additional message processing overheads in terms of number of items searched during demultiplexing and amount of time spent searching because it has the potential to increase the number of messages exchanged and to introduce non-deterministic message ordering. Therefore, understanding the implications of adding multithreading to MPI applications is important for future application development.One strategy for advancing this understanding is through ‘low-cost’ benchmarks that emulate full communication patterns using fewer resources. For example, while a complete, ‘real-world’ multithreaded halo exchange requires 9 or 27 nodes, the low-cost alternative needs only two, making it deployable on systems where acquiring resources is difficult because of high utilization (e.g., busy capacity-computing systems), or impossible because the necessary resources do not exist (e.g., testbeds with too few nodes). While such benchmarks have been proposed, the reported results have been limited to a single architecture or derived indirectly through simulation, and no attempt has been made to confirm that a low-cost benchmark accurately captures features of full (non-emulated) exchanges. Moreover, benchmark code has not been made publicly available.The purpose of the study presented in this paper is to quantify how accurately the low-cost benchmark captures the matching behavior of the full, real-world benchmark. In the process, we also advocate for the feasibility and utility of the low-cost benchmark. We present a ‘real-world’ benchmark implementing a full multithreaded halo exchange on 9 and 27 nodes, as defined by 5-point and 9-point 2D stencils, and 7-point and 27-point 3D stencils. Likewise, we present a ‘low-cost’ benchmark that emulates these communication patterns using only two nodes. We then confirm, across multiple architectures, that the low-cost benchmark gives accurate estimates of both number of items searched during message processing, and time spent processing those messages. Finally, we demonstrate the utility of the low-cost benchmark by using it to profile the performance impact of state-of-the-art Mellanox ConnectX-5 hardware support for offloaded MPI message demultiplexing. To facilitate further research on the effects of multithreaded MPI on message matching behavior, the source of our two benchmarks is to be included in the next release version of the Sandia MPI Micro-Benchmark Suite.
给 MPI 线程一个公平的机会:多线程 MPI 设计研究
DOI: 10.1109/cluster.2019.8891015
发表时间: 2019
期刊: IEEE Cluster
影响因子: --
作者:
Patinyasakdikul, T.;Eberius, D.;Bosilca, G.;Hjelm, N.
通讯作者: Hjelm, N.