DPU-Bench: A Micro-Benchmark Suite to Measure Offload Efficiency Of SmartNICs

DPU-Bench: A Micro-Benchmark Suite to Measure Offload Efficiency Of SmartNICs
复制标题

DPU-Bench:用于测量 SmartNIC 卸载效率的微基准套件

DOI:
10.1145/3569951.3593595
复制
发表时间:
2023
期刊:
Practice and Experience in Advanced Research Computing 23
影响因子:
--
通讯作者:
Poole, Steve
Poole, Steve
中科院分区:
--
文献类型:
--
作者:
Michalowicz, Benjamin;Kandadi Suresh, Kaushik;Subramoni, Hari;Panda, Dhabaleswar;Poole, Steve

文献摘要

参考文献

被引文献

相似文献

智能网络接口卡 (SmartNIC) 在过去几年中的受欢迎程度大幅增长,例如 NVIDIA BlueField-2 数据处理单元 (DPU)。配备自己的一组核心和内存使它们能够执行常规 NIC 之外的操作,HPC 研究人员正在设计使用它们的新方法。例如,将通信卸载到一个CPU“主机”可以执行计算量更大的任务。然而,仍然存在一个问题:在面临性能下降之前,有多少工作可以分配给 SmartNIC 上的进程?我们推出了 DPU-Bench:一种使用 IB-Verbs 原语的低级微基准套件,使 HPC 用户能够检查要放置在一个或多个 SmartNIC 上的进程数量,以便有效地卸载给定的通信模式。我们在本文中以不同的工作分配机制检查了中等规模的直接算法,并深入了解了不同数量的工作进程和消息大小所发现的趋势。
Smart Network Interface Cards (SmartNIC) have experienced massive growth in popularity over the last few years such as the NVIDIA BlueField-2 Data Processing Unit (DPU). Being equipped with their own set of cores and memory allows them to perform actions beyond a regular NIC, and HPC researchers are designing new ways to use them. For example, offloading communication to one enables the CPU "host" to perform more computationally heavy tasks. However, one question remains: How much of that work can be distributed among processes placed on the SmartNIC before facing performance degradation? We present DPU-Bench: A low-level micro-benchmark suite using IB-Verbs primitives to enable HPC users to examine the number of processes to be placed on one or more SmartNICs in order to efficiently offload a given communication pattern. We examine direct algorithms in this paper at a medium scale with different work assignment mechanisms and give insights into the trends found with varying numbers of worker processes and message sizes.
DOI: 10.1007/978-3-030-78713-4_2
发表时间: 2021
期刊: --
影响因子: --
作者:
Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda
通讯作者: Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda
INAM:基于 MPI 的应用程序中通信的跨堆栈分析和分析
DOI: 10.1145/3437359.3465582
发表时间: 2021
期刊: Practice and Experience in Advanced Research Computing
影响因子: --
作者:
Pouya Kousha;Kamal Raj Sankarapandian Dayala Ganesh Ram;M. Kedia;H. Subramoni;Arpan Jain;A. Shafi;D. Panda;Trey Dockendorf;Heechang Na;K. Tomko
通讯作者: K. Tomko
通过 BlueField-2 DPU 实现大消息非阻塞 MPI_Iallgather 和 MPI Ibcast 卸载
DOI: 10.1109/hipc53243.2021.00054
发表时间: 2021
期刊: and Analytics (HiPC
影响因子: --
作者:
Sarkauskas, Nick;Bayatpour, Mohammadreza;Tran, Tu;Ramesh, Bharath;Subramoni, Hari;Panda, Dhabaleswar K.
通讯作者: Panda, Dhabaleswar K.
使用 BlueField-2 DPU 加速现代 HPC 集群上基于 CPU 的分布式 DNN 训练
DOI: 10.1109/hoti52880.2021.00017
发表时间: 2021
期刊: 2021 IEEE Symposium on High-Performance Interconnects (HOTI
影响因子: --
作者:
Jain, Arpan;Alnaasan, Nawras;Shafi, Aamir;Subramoni, Hari;Panda, Dhabaleswar K
通讯作者: Panda, Dhabaleswar K
Nvidia 数据中心处理单元 (DPU) 架构
DOI: 10.1109/hcs52781.2021.9567066
发表时间: 2021
期刊: 2021 IEEE Hot Chips 33 Symposium (HCS)
影响因子: --
作者:
Idan Burstein
通讯作者: Idan Burstein