Reconfigurable switches for high performance and flexible MPI collectives

Reconfigurable switches for high performance and flexible MPI collectives
复制标题

DOI:
10.1002/cpe.6769
复制
发表时间:
2021-12
期刊:
Concurrency and Computation: Practice and Experience
影响因子:
--
通讯作者:
Pouya Haghi;Anqi Guo;Qingqing Xiong;Chen Yang;Tong Geng;Justin T. Broaddus;Ryan J. Marshall;Derek Schafer;A. Skjellum;Martin C. Herbordt
Pouya Haghi;Anqi Guo;Qingqing Xiong;Chen Yang;Tong Geng;Justin T. Broaddus;Ryan J. Marshall;Derek Schafer;A. Skjellum;Martin C. Herbordt
中科院分区:
其他
文献类型:
--
作者:
Pouya Haghi;Anqi Guo;Qingqing Xiong;Chen Yang;Tong Geng;Justin T. Broaddus;Ryan J. Marshall;Derek Schafer;A. Skjellum;Martin C. Herbordt

文献摘要

被引文献

相似文献

在将MPI集合操作转移到硬件方面已经做出了很多努力。但是,尽管基于NIC的集体加速已经得到了很好的研究,但将它们的处理负载转移到交换结构中,尽管有许多优势,但却要有限得多。固定逻辑实现的一个主要问题是,要么只加速了可能的集体通信的一小部分,要么在不需要特定功能的应用程序中浪费了逻辑。使用可重构逻辑有许多优点:可以准确地实现所需的操作;可以指定所需性能的级别;可以定义和实现新的、可能复杂的操作。我们已经设计了一个交换机内集合加速器MPI-FPGA,并通过七个MPI集合以及一组基准测试和代理应用程序(MiniApp)演示了它的使用。该加速器采用了一种新颖的包含全流水线矢量化聚合逻辑单元的两级交换设计。这项工作的关键是对子通信器集体提供支持,以支持任意形状的通信器,并且可扩展到大型系统。流接口提高了长消息的性能。虽然这种可重构设计普遍适用,但我们使用了以FPGA为中心的集群作为原型。在最有可能的情况下,直接网络中的样本MPI-FPGA设计比传统集群实现了相当大的加速比。我们还给出了具有可重构高基交换机的间接网络的结果,并表明该方法在夏普支持的操作子集方面与夏普技术具有竞争性。MPI-FPGA完全集成到MPICH中,对MPI应用程序是透明的。
There has been much effort in offloading MPI collective operations into hardware. But while NIC‐based collective acceleration is well‐studied, offloading their processing into the switching fabric, despite numerous advantages, has been much more limited. A major problem with fixed logic implementations is that either only a fraction of the possible collective communication is accelerated or that logic is wasted in the applications that do not need a particular capability. Using reconfigurable logic has numerous advantages: exactly the required operations can be implemented; the level of desired performance can be specified; and new, possibly complex, operations can be defined and implemented. We have designed an in‐switch collective accelerator, MPI‐FPGA, and demonstrated its use with seven MPI collectives and over a set of benchmarks and proxy applications (MiniApps). The accelerator uses a novel two‐level switch design containing fully pipelined vectorized aggregation logic units. Essential to this work is providing support for sub‐communicator collectives that enables communicators of arbitrary shape, and that is scalable to large systems. A streaming interface improves the performance for long messages. While this reconfigurable design is generally applicable, we prototype it with an FPGA‐centric cluster. A sample MPI‐FPGA design in a direct network achieves considerable speedups over conventional clusters in the most likely scenarios. We also present results for indirect networks with reconfigurable high‐radix switches and show that this approach is competitive with SHArP technology for the subset of operations that SHArP supports. MPI‐FPGA is fully integrated into MPICH and is transparent to MPI applications.