How I learned to stop worrying about user-visible endpoints and love MPI

How I learned to stop worrying about user-visible endpoints and love MPI
复制标题

我如何学会不再担心用户可见的端点并爱上 MPI

DOI:
10.1145/3392717.3392773
复制
发表时间:
2020
期刊:
Proceedings of the 34th ACM International Conference on Supercomputing
影响因子:
--
通讯作者:
P. Balaji
P. Balaji
中科院分区:
--
文献类型:
--
作者:
Rohit Zambre;Aparna Chandramowlishwaran;P. Balaji

文献摘要

参考文献

被引文献

相似文献

MPI+线程作为传统的“MPI无处不在”模型的替代方案越来越突出,以便更好地处理与其他节点上资源相比核数量不成比例的增加。但是,MPI+线程的通信性能可能比MPI的任何地方都慢100倍。MPI用户和开发人员都应该为这种放缓负责。MPI用户传统上没有公开逻辑通信并行性。因此,MPI库使用保守的方法,如全局临界区,来维护MPI+线程的MPI排序约束,从而串行化对底层并行网络资源的访问并限制性能。为了增强MPI+线程的通信性能,研究人员提出了MPI端点作为MPI-3.1标准的用户可见扩展。MPI端点允许单个进程在通信器内创建多个MPI等级。从理论上讲,这可以允许每个线程都有一个专用的网络通信路径,从而避免线程之间的资源争用并提高性能。然而,将线程映射到端点的责任将由领域科学家承担。在本文中,我们扮演魔鬼的拥护者的角色,并质疑这种用户可见的端点的必要性。我们当然同意,专用的沟通渠道至关重要。然而,在多大程度上,我们可以在不修改MPI标准的情况下将这些通道隐藏在MPI库中,从而减轻用户的负担?更重要的是,通过这种抽象,我们会失去什么功能?本文通过MPI-3.1标准的新实现来回答这些问题,该标准在MPI库中使用多个虚拟通信接口(VCI)。VCI抽象底层网络上下文。当用户通过现有的MPI机制公开并行性时,MPI库将该并行性映射到VCI,从而使领域科学家不必担心端点。我们确定的情况下,用户暴露的并行VC执行以及用户可见的端点,以及这种抽象损害性能的情况下。
MPI+threads is gaining prominence as an alternative to the traditional "MPI everywhere" model in order to better handle the disproportionate increase in the number of cores compared with other on-node resources. However, the communication performance of MPI+threads can be 100x slower than that of MPI everywhere. Both MPI users and developers are to blame for this slowdown. MPI users traditionally have not exposed logical communication parallelism. Consequently, MPI libraries have used conservative approaches, such as a global critical section, to maintain MPI's ordering constraints for MPI+threads, thus serializing access to the underlying parallel network resources and limiting performance. To enhance the communication performance of MPI+threads, researchers have proposed MPI Endpoints as a user-visible extension to the MPI-3.1 standard. MPI Endpoints allows a single process to create multiple MPI ranks within a communicator. This could, in theory, allow each thread to have a dedicated communication path to the network, thus avoiding resource contention between threads and improving performance. The onus of mapping threads to endpoints, however, would then be on domain scientists. In this paper we play the role of devil's advocate and question the need for such user-visible endpoints. We certainly agree that dedicated communication channels are critical. To what extent, however, can we hide these channels inside the MPI library without modifying the MPI standard and thus unburden the user? More important, what functionality would we lose through such abstraction? This paper answers these questions through a new implementation of the MPI-3.1 standard that uses multiple virtual communication interfaces (VCIs) inside the MPI library. VCIs abstract underlying network contexts. When users expose parallelism through existing MPI mechanisms, the MPI library maps that parallelism to the VCIs, relieving the domain scientists from worrying about endpoints. We identify cases where user-exposed parallelism on VCIs perform as well as user-visible endpoints, as well as cases where such abstraction hurts performance.
给 MPI 线程一个公平的机会:多线程 MPI 设计研究
DOI: 10.1109/cluster.2019.8891015
发表时间: 2019
期刊: IEEE Cluster
影响因子: --
作者:
Patinyasakdikul, T.;Eberius, D.;Bosilca, G.;Hjelm, N.
通讯作者: Hjelm, N.