FLASH: FPGA-Accelerated Smart Switches with GCN Case Study
FLASH: FPGA-Accelerated Smart Switches with GCN Case Study
复制标题
FLASH:采用 GCN 的 FPGA 加速智能开关案例研究
DOI:
10.1145/3577193.3593739
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Li, Ang
中科院分区:
文献类型:
--
作者:
Haghi, Pouya;Krska, William;Tan, Cheng;Geng, Tong;Chen, Po Hao;Greenwood, Connor;Guo, Anqi;Hines, Thomas;Wu, Chunshu;Li, Ang
Some communication switches, e.g., the Mellanox SHArP and those in the IBM BlueGene clusters, are augmented to process packets at the application level with fixed-function collectives. This approach, however, lacks flexibility, which limits their applicability in diverse and dynamic workloads. Recently, a new type of programmable packet processor, which uses high-level languages,e.g., P4, has emerged as a possible candidate. P4-based switches, however, fall short in certain applications, including machine learning, where capabilities not currently supported by P4 are needed. These include more complex calculation, such as sparse computation and fused multiply-accumulate, data-intensive floating point operations, data reuse, and significant memory. The problem addressed here is that such a switch augmentation needs to support: a large amount of state, significant flexible compute capability, and ease of programming, all while maintaining full functionality, including ensuring high throughput, and demonstrating utility.In this work, we propose a programmable look-aside-type accelerator that can be embedded into, or attached to, existing communication switch pipelines and that is capable of processing packets at line rate. The proposed in-switch accelerator is based on mixing an ISA (subset of RISC-V instructions) with dataflow graphs (found in CGRAs). To augment performance, vector instructions are also supported. To facilitate usability, we have developed a complete toolchain to compile user-provided C/C++ codes to appropriate back-end instructions for configuring the accelerator. While this approach is flexible enough to support various workloads, in this paper, we consider Graph Convolutional Networks (GCNs) as a case study. Experimental results show that this approach considerably improves the performance of distributed GCN applications.
登录
查看更多内容
DOI:
--
发表时间:
2019
期刊:
IEEE Symposium on Field-Programmable Custom Computing Machines
影响因子:
--
作者:
Ian Taras;J. Anderson
通讯作者:
J. Anderson
DOI:
--
发表时间:
2019
期刊:
影响因子:
--
作者:
Joshua Stern;Qingqing Xiong;A. Skjellum
通讯作者:
A. Skjellum
DOI:
10.1007/978-3-030-50743-5_3
发表时间:
2020-05-22
期刊:
High Performance Computing
影响因子:
--
作者:
Graham RL;Levi L;Burredy D;Bloch G;Shainer G;Cho D;Elias G;Klein D;Ladd J;Maor O;Marelli A;Petrov V;Romlet E;Qin Y;Zemah I
通讯作者:
Zemah I
DOI:
--
发表时间:
2020
期刊:
IEEE Symposium on High-Performance Interconnects
影响因子:
--
作者:
V. Krishnan;O. Serres;M. Blocksome
通讯作者:
M. Blocksome
DOI:
10.1007/978-3-030-78713-4_2
发表时间:
2021
期刊:
--
影响因子:
--
作者:
Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda
通讯作者:
Mohammadreza Bayatpour;Nick Sarkauskas;H. Subramoni;J. Hashmi;D. Panda