Scalable Communication Endpoints for MPI+Threads Applications

Scalable Communication Endpoints for MPI+Threads Applications
复制标题

MPI 线程应用程序的可扩展通信端点

DOI:
10.1109/padsw.2018.8645059
复制
发表时间:
2018
期刊:
2018 IEEE 24th International Conference on Parallel and Distributed Systems (ICPADS)
影响因子:
--
通讯作者:
P. Balaji
P. Balaji
中科院分区:
--
文献类型:
--
作者:
Rohit Zambre;Aparna Chandramowlishwaran;P. Balaji

文献摘要

被引文献

相似文献

混合MPI+线程编程作为传统的“MPI无处不在”模型的替代方案越来越突出,以更好地处理与其他节点上资源相比核数量不成比例的增加。这两个模型的当前实现代表了现代MPI实现中通信资源共享的两种极端情况。在MPI-everywhere模型中,每个MPI进程都有一组专用的通信资源(也称为端点),这对性能来说是理想的,但浪费资源。使用MPI+线程,当前的MPI实现为所有线程共享一个通信端点,这对于资源使用是理想的,但对性能是有害的。在本文中,我们探讨了MPI+线程环境中的性能和通信资源使用之间的权衡空间。我们首先演示了两种极端情况,一种是所有线程共享一个通信端点,另一种是每个线程都有自己的专用通信端点(类似于MPI无处不在模型),并展示了这两种情况下的效率低下。接下来,我们将对Mellanox InfiniBand环境中的不同级别的资源共享进行全面分析。利用从这种分析中吸取的经验教训,我们设计了一个改进的资源共享模型,以产生可扩展的通信端点,可以实现相同的性能,每个线程的专用通信资源,但只使用三分之一的资源。
Hybrid MPI+threads programming is gaining prominence as an alternative to the traditional “MPI everywhere” model to better handle the disproportionate increase in the number of cores compared with other on-node resources. Current implementations of these two models represent the two extreme cases of communication resource sharing in modern MPI implementations. In the MPI-everywhere model, each MPI process has a dedicated set of communication resources (also known as endpoints), which is ideal for performance but is resource wasteful. With MPI+threads, current MPI implementations share a single communication endpoint for all threads, which is ideal for resource usage but is hurtful for performance. In this paper, we explore the tradeoff space between performance and communication resource usage in MPI+threads environments. We first demonstrate the two extreme cases-one where all threads share a single communication endpoint and another where each thread gets its own dedicated communication endpoint (similar to the MPI-everywhere model) and showcase the inefficiencies in both these cases. Next, we perform a thorough analysis of the different levels of resource sharing in the context of Mellanox InfiniBand. Using the lessons learned from this analysis, we design an improved resource-sharing model to produce scalable communication endpoints that can achieve the same performance as with dedicated communication resources per thread but using just a third of the resources.