Hardware implementation of MPI_Barrier on an FPGA cluster

Hardware implementation of MPI_Barrier on an FPGA cluster
复制标题

MPI_Barrier 在 FPGA 集群上的硬件实现

DOI:
10.1109/fpl.2009.5272560
复制
发表时间:
2009
期刊:
2009 International Conference on Field Programmable Logic and Applications
影响因子:
--
通讯作者:
R. Sass
R. Sass
中科院分区:
--
文献类型:
--
作者:
Shanyuan Gao;A. Schmidt;R. Sass

文献摘要

被引文献

相似文献

消息传递是分布存储并行计算机的主要编程模型,消息传递接口(MPI)是其标准。沿着点对点发送和接收消息原语,MPI包括一组用于同步和协调任务组的集体通信操作。MPI_Barrier是最重要的集合过程之一,在过去的二十年中,已经在各种架构上进行了广泛的研究。然而,一组平台FPGA是一种新的架构,为实现屏障操作提供了有趣的、资源高效的选择。介绍了MPI Barrier的FPGA实现。前提是,屏障(和其他集体通信操作)对延迟非常敏感,因为节点的数量可以扩展到数万个。FPGA上相对较慢的处理器将显著限制性能。FPGA硬件设计实现了基于树的算法,并与定制的高速片上/片外网络紧密集成。MPI访问通过专门设计的内核模块提供。这有效地将工作从CPU和操作系统卸载到硬件中。该设计的评估结果表明,与传统的FPGA集群和商品集群上的软件实现相比,性能显著提高。此外,它表明将其他MPI集体操作转移到硬件中将是有益的。
Message-Passing is the dominant programming model for distributed memory parallel computers and Message- Passing Interface (MPI) is the standard. Along with pointto- point send and receive message primitives, MPI includes a set of collective communication operations that are used to synchronize and coordinate groups of tasks. The MPI_Barrier, one of the most important collective procedures, has been extensively studied on a variety of architectures over last twenty years. However, a cluster of Platform FPGAs is a new architecture and offers interesting, resourceefficient options for implementing the barrier operation. This paper describes an FPGA implementation of MPI Barrier. The premise is that barrier (and other collective communication operations) are very sensitive to latency as the number of nodes scales to the tens-of-thousands. The relatively slow processors found on FPGAs will significantly cap performance. The FPGA hardware design implements a tree-based algorithm and is tightly integrated with the custom high-speed on-chip/off-chip network. MPI access is available through a specially-designed kernel module. This effectively offloads the work from the CPU and OS into hardware. The evaluation of this design shows signficant performance gains compared with a conventional software implementation on both an FPGA cluster and a commodity cluster. Further, it suggests that moving other MPI collective operations into hardware would be beneficial.