FLASH: FPGA-Accelerated Smart Switches with GCN Case Study

FLASH: FPGA-Accelerated Smart Switches with GCN Case Study
复制标题

FLASH:采用 GCN 的 FPGA 加速智能开关案例研究

DOI:
10.1145/3577193.3593739
复制
发表时间:
2023
期刊:
ICS 2023: International Conference on Supercomputing
影响因子:
--
通讯作者:
Li, Ang
Li, Ang
中科院分区:
--
文献类型:
--
作者:
Haghi, Pouya;Krska, William;Tan, Cheng;Geng, Tong;Chen, Po Hao;Greenwood, Connor;Guo, Anqi;Hines, Thomas;Wu, Chunshu;Li, Ang

文献摘要

参考文献

被引文献

相似文献

一些通信交换机,例如,MellanoxSHArP和IBMBlueGene集群中的那些被增强以在具有固定功能集合的应用程序级处理分组。然而,这种方法缺乏灵活性,这限制了它们在多样化和动态工作负载中的适用性。最近,一种新型的可编程分组处理器,其使用高级语言,例如,P4已经成为可能的候选人。然而,基于P4的交换机在某些应用中存在不足,包括机器学习,其中需要P4当前不支持的功能。这些包括更复杂的计算,如稀疏计算和融合乘法累加,数据密集型浮点运算,数据重用和大量内存。这里解决的问题是,这种交换机增强需要支持:大量的状态,显着的灵活计算能力,易于编程,同时保持完整的功能,包括确保高吞吐量,并演示utility.In这项工作中,我们提出了一个可编程的后备型加速器,可以嵌入到,或连接到,现有的通信交换流水线,并且能够以线路速率处理分组。提议的交换机内加速器基于混合伊萨(RISC-V指令的子集)和CRAs中的CRAs图。为了增强性能,还支持向量指令。为了提高可用性,我们开发了一个完整的工具链,将用户提供的C/C++代码编译为适当的后端指令,以配置加速器。虽然这种方法足够灵活,可以支持各种工作负载,但在本文中,我们将图卷积网络(GCN)作为案例研究。实验结果表明,该方法大大提高了分布式GCN应用程序的性能。
Some communication switches, e.g., the Mellanox SHArP and those in the IBM BlueGene clusters, are augmented to process packets at the application level with fixed-function collectives. This approach, however, lacks flexibility, which limits their applicability in diverse and dynamic workloads. Recently, a new type of programmable packet processor, which uses high-level languages,e.g., P4, has emerged as a possible candidate. P4-based switches, however, fall short in certain applications, including machine learning, where capabilities not currently supported by P4 are needed. These include more complex calculation, such as sparse computation and fused multiply-accumulate, data-intensive floating point operations, data reuse, and significant memory. The problem addressed here is that such a switch augmentation needs to support: a large amount of state, significant flexible compute capability, and ease of programming, all while maintaining full functionality, including ensuring high throughput, and demonstrating utility.In this work, we propose a programmable look-aside-type accelerator that can be embedded into, or attached to, existing communication switch pipelines and that is capable of processing packets at line rate. The proposed in-switch accelerator is based on mixing an ISA (subset of RISC-V instructions) with dataflow graphs (found in CGRAs). To augment performance, vector instructions are also supported. To facilitate usability, we have developed a complete toolchain to compile user-provided C/C++ codes to appropriate back-end instructions for configuring the accelerator. While this approach is flexible enough to support various workloads, in this paper, we consider Graph Convolutional Networks (GCNs) as a case study. Experimental results show that this approach considerably improves the performance of distributed GCN applications.
FPGA 架构对 CGRA 覆盖面积和性能的影响
DOI: --
发表时间: 2019
期刊: IEEE Symposium on Field-Programmable Custom Computing Machines
影响因子: --
作者:
Ian Taras;J. Anderson
通讯作者: J. Anderson
支持通信器进行 MPI 集合的交换内处理的新方法
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者:
Joshua Stern;Qingqing Xiong;A. Skjellum
通讯作者: A. Skjellum
DOI: 10.1007/978-3-030-50743-5_3
发表时间: 2020-05-22
期刊: High Performance Computing
影响因子: --
作者:
Graham RL;Levi L;Burredy D;Bloch G;Shainer G;Cho D;Elias G;Klein D;Ladd J;Maor O;Marelli A;Petrov V;Romlet E;Qin Y;Zemah I
通讯作者: Zemah I
COConfigurable 网络协议加速器 (COPA) †:集成网络/加速器硬件/软件框架
DOI: --
发表时间: 2020
期刊: IEEE Symposium on High-Performance Interconnects
影响因子: --
作者:
V. Krishnan;O. Serres;M. Blocksome
通讯作者: M. Blocksome
DOI: 10.1007/978-3-030-78713-4_2
发表时间: 2021
期刊: --
影响因子: --
作者:
Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda
通讯作者: Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda