BluesMPI: Efficient MPI Non-blocking Alltoall Offloading Designs on Modern BlueField Smart NICs

BluesMPI: Efficient MPI Non-blocking Alltoall Offloading Designs on Modern BlueField Smart NICs
复制标题

DOI:
10.1007/978-3-030-78713-4_2
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda
Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda
中科院分区:
其他
文献类型:
--
作者:
Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda

文献摘要

被引文献

相似文献

在最先进的生产质量MPI(消息传递接口)库中,通信进程由主线程或单独的通信进程线程执行。利用单独的通信线程可以导致通信和计算的更高重叠,以及减少总的应用程序执行时间。然而,这样的方法也可能导致CPU资源的争用,从而导致低于标准的应用性能,因为应用本身具有较少数量的可用于计算的核。最近,Mellanox推出了BlueField系列适配器,它将传统的基于ASIC的网络适配器的先进功能与ARM处理器阵列相结合。在本文中,我们提出了BluesMPI,一个高性能的MPI非阻塞Alltoall设计,可用于卸载MPI_Ialltoall集体操作从主机CPU到智能网卡。BluesMPI保证了Alltoall集体操作的通信和计算的完全重叠,同时为基于CPU的加载设计提供了同等的纯通信延迟。我们探讨了几种设计,以实现最佳的纯通信延迟MPI_Ialltoall。我们的实验表明,BluesMPI可以提高总的执行时间的OSU微基准MPI_Ialltoall和P3DFFT应用程序高达44%和30%,分别。据我们所知,这是第一个有效利用现代BlueField Smart的设计,用于派生MPI Alltoall集体操作,以获得通信和计算的峰值重叠。
In the state-of-the-art production quality MPI (Message Passing Interface) libraries, communication progress is either performed by the main thread or a separate communication progress thread. Taking advantage of separate communication threads can lead to a higher overlap of communication and computation as well as reduced total application execution time. However, such an approach can also lead to contention for CPU resources leading to sub-par application performance as the application itself has less number of available cores for computation. Recently, Mellanox has introduced the BlueField series of adapters which combine the advanced capabilities of traditional ASIC based network adapters with an array of ARM processors. In this paper, we propose BluesMPI, a high performance MPI non-blocking Alltoall design that can be used to offload MPI_Ialltoall collective operations from the host CPU to the Smart NIC. BluesMPI guarantees the full overlap of communication and computation for Alltoall collective operations while providing on-par pure communication latency to CPU based on-loading designs. We explore several designs to achieve the best pure communication latency for MPI_Ialltoall. Our experiments show that BluesMPI can improve the total execution time of the OSU Micro Benchmark for MPI_Ialltoall and P3DFFT application up to 44% and 30%, respectively. To the best of our knowledge, this is the first design that efficiently takes advantage of modern BlueField Smart NICs in deriving the MPI Alltoall collective operation to get peak overlap of communication and computation.