Scalable Hierarchical Aggregation and Reduction Protocol (SHARP)TM Streaming-Aggregation Hardware Design and Evaluation

Scalable Hierarchical Aggregation and Reduction Protocol (SHARP)TM Streaming-Aggregation Hardware Design and Evaluation
复制标题

DOI:
10.1007/978-3-030-50743-5_3
复制
发表时间:
2020-05-22
期刊:
High Performance Computing
影响因子:
--
通讯作者:
Zemah I
Zemah I
中科院分区:
其他
文献类型:
--
作者:
Graham RL;Levi L;Burredy D;Bloch G;Shainer G;Cho D;Elias G;Klein D;Ladd J;Maor O;Marelli A;Petrov V;Romlet E;Qin Y;Zemah I

文献摘要

被引文献

相似文献

本文介绍了新的基于硬件的流聚合功能添加到Mellanox的可扩展分层聚合和减少协议在其HDR InfiniBand交换机。对于大型消息,此功能旨在实现类似于相同大小的点对点消息的缩减带宽,并补充延迟优化的低延迟聚合缩减功能,旨在减少小数据。在基于HDR InfiniBand的系统上测量的MPI_Allreduce()带宽达到网络带宽的约95%。对于中型和大型数据缩减,这也将缩减带宽相对于基于主机的(例如,基于软件的)归约算法。使用此功能还将DL-Poly和PyTorch应用程序的性能分别提高了4%和18%。本文介绍了SHARP流聚合硬件架构和一组合成和应用基准用于研究这种新的减少能力,以及流聚合的数据大小的范围比低延迟聚合算法表现更好。
This paper describes the new hardware-based streaming-aggregation capability added to Mellanox’s Scalable Hierarchical Aggregation and Reduction Protocol in its HDR InfiniBand switches. For large messages, this capability is designed to achieve reduction bandwidths similar to those of point-to-point messages of the same size, and complements the latency-optimized low-latency aggregation reduction capabilities, aimed at small data reductions. MPI_Allreduce() bandwidth measured on an HDR InfiniBand based system achieves about 95% of network bandwidth. For medium and large data reduction this also improves the reduction bandwidth by a factor of 2–5 relative to host-based (e.g., software-based) reduction algorithms. Using this capability also increased DL-Poly and PyTorch application performance by as much as 4% and 18%, respectively. This paper describes SHARP Streaming-Aggregation hardware architecture and a set of synthetic and application benchmarks used to study this new reduction capability, and the range of data sizes for which Streaming-Aggregation performs better than the low-latency aggregation algorithm.