A Novel Approach to Supporting Communicators for In-Switch Processing of MPI Collectives
A Novel Approach to Supporting Communicators for In-Switch Processing of MPI Collectives
复制标题
支持通信器进行 MPI 集合的交换内处理的新方法
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
A. Skjellum
中科院分区:
文献类型:
--
作者:
Joshua Stern;Qingqing Xiong;A. Skjellum
MPI collective operations can often be performance killers in HPC applications; we seek to solve this bottleneck by offloading them to hardware within the switch itself. We have seen from previous work that moving collectives into the network offers significant performance benefits. However, there has been little advancement in providing support for sub-communicator collectives. We introduce a novel mechanism that enables the hardware to support a large number of communicators of arbitrary shape that is scalable to very large systems. We have integrated this support into an in-switch hardware accelerator to implement support for MPI communicators and full offload of MPI collectives. While this mechanism is universally applicable, we implement it in an FPGA cluster; FPGAs provide the ability to couple communication and computation and so provide an ideal testbed. Preliminary results show that we can achieve substantial performance improvement at acceptable hardware cost, including a 10× speedup over conventional clusters for short message collectives over irregular intra-communicators.