A high performance broadcast design with hardware multicast and GPUDirect RDMA for streaming applications on Infiniband clusters

A high performance broadcast design with hardware multicast and GPUDirect RDMA for streaming applications on Infiniband clusters
复制标题

具有硬件多播和 GPUDirect RDMA 的高性能广播设计,适用于 Infiniband 集群上的流应用程序

DOI:
10.1109/hipc.2014.7116875
复制
发表时间:
2014
期刊:
2014 21st International Conference on High Performance Computing (HiPC)
影响因子:
--
通讯作者:
D. Panda
D. Panda
中科院分区:
--
文献类型:
--
作者:
Akshay Venkatesh;H. Subramoni;Khaled Hamidouche;D. Panda

文献摘要

参考文献

被引文献

相似文献

在高性能计算领域的几个流应用程序通过利用现代GPGPU提供的原始计算能力在执行时间上获得显著的加速。这种原始计算能力,加上高性能互连(如InfiniBand(IB))提供的高网络吞吐量,使流媒体应用能够快速扩展。构成多节点流应用的执行的频繁使用的操作是广播操作,其中来自单个源的数据被发送到多个接收器,通常是从实况数据站点。虽然像IB这样的高性能网络提供了新的功能,如基于硬件的多播,以加快广播操作的性能,但由于IB主机通道适配器(HCA)无法直接访问GPGPU的存储器,因此它们的好处仅限于基于主机的应用程序。这对严重依赖来自GPU存储器的广播操作的高性能流应用造成了显著的性能瓶颈。最近推出的GPUDirect RDMA功能通过使IB HCA能够直接执行与GPU内存之间的数据传输(绕过主机内存)来消除这一瓶颈。因此,它提出了一个有吸引力的替代设计高性能的广播操作的GPGPU为基础的高性能流应用程序。在这项工作中,我们提出了一种新的方法,充分利用GPUDirect的RDMA和硬件多播功能串联设计一个高性能的流媒体应用程序的广播操作。与64个GPU节点上的朴素方案相比,使用所提出的设计进行的实验显示延迟降低了60%,吞吐量基准提高了3倍至4倍。
Several streaming applications in the field of high performance computing are obtaining significant speedups in execution time by leveraging the raw compute power offered by modern GPGPUs. This raw compute power, coupled with the high network throughput offered by high performance interconnects such as InfiniBand (IB) are allowing streaming applications to scale to rapidly. A frequently used operation that constitutes to the execution of multi-node streaming applications is the broadcast operation where data from a single source is transmitted to multiple sinks, typically from a live data site. Although high performance networks like IB offer novel features like hardware based multicast to speed up the performance of the broadcast operation, their benefits have been limited to host based applications due to the inability of IB Host Channel Adapters (HCAs) to directly access the memory of the GPGPUs. This poses a significant performance bottleneck to high performance streaming applications that rely heavily on broadcast operations from GPU memories. The recently introduced GPUDirect RDMA feature alleviates this bottleneck by enabling IB HCAs to perform data transfers directly to / from GPU memory (bypassing host memory). Thus, it presents an attractive alternative to designing high performance broadcast operations for GPGPU based high performance streaming applications. In this work, we propose a novel method for fully utilizing GPUDirect RDMA and hardware multicast features in tandem to design a high performance broadcast operation for streaming applications. The experiments conducted with the proposed design show up 60% decrease in latency and 3X-4X improvement in a throughput benchmark compared to the naive scheme on 64 GPU nodes.
DOI: 10.1177/1094342010391989
发表时间: 2011-02-01
影响因子: 3.1
作者:
Dongarra, Jack;Beckman, Pete;Yelick, Kathy
通讯作者: Yelick, Kathy