A Framework for Neural Network Inference on FPGA-Centric SmartNICs

A Framework for Neural Network Inference on FPGA-Centric SmartNICs
复制标题

DOI:
10.1109/fpl57034.2022.00071
复制
发表时间:
2022-08
期刊:
2022 32nd International Conference on Field-Programmable Logic and Applications (FPL)
影响因子:
--
通讯作者:
Anqi Guo;Tong Geng;Yongan Zhang;Pouya Haghi;Chunshu Wu;Cheng Tan;Yingyan Lin;Ang Li;Martin C. Herbordt
Anqi Guo;Tong Geng;Yongan Zhang;Pouya Haghi;Chunshu Wu;Cheng Tan;Yingyan Lin;Ang Li;Martin C. Herbordt
中科院分区:
其他
文献类型:
--
作者:
Anqi Guo;Tong Geng;Yongan Zhang;Pouya Haghi;Chunshu Wu;Cheng Tan;Yingyan Lin;Ang Li;Martin C. Herbordt

文献摘要

被引文献

相似文献

基于FPGA的SmartCable提供了巨大的潜力,通过紧密耦合支持可重构数据密集型计算与跨节点通信,从而缓解冯诺依曼瓶颈,显着提高高性能计算和仓库数据处理的性能。然而,现有的工作通常是有限的,因为它假设一个加速器模型,其中内核被卸载到智能CPU与大多数控制任务留给CPU。这会导致频繁的等待,降低性能和扩展挑战。在这项工作中,我们提出了一个新的分布式数据为中心的计算框架FCSN可重构的SmartNIC为基础的系统。FCsN通过一个轻量级的任务循环执行模型及其实现架构,使得神经网络内核的执行控制逻辑、系统调度和网络通信完全脱离了智能机。这通过(i)避免控制对CPU的依赖以及(ii)以线速率和非常细粒度的方式支持流式NN内核执行和网络通信来提高性能。我们使用各种类型的神经网络内核和应用程序(包括深度神经网络(DNN)和图形神经网络(GNN))来展示FCsN的效率和灵活性,因为后者既不规则又是数据密集型的,它们提供了特别强大的演示。使用常用的神经网络模型和图形数据集进行的评估表明,与基于MPI的标准CPU基线相比,具有FCsN的系统可以实现10倍的加速比
FPGA-based SmartNICs offer great potential to significantly improve the performance of high-performance computing and warehouse data processing by tightly coupling support for reconfigurable data-intensive computation with cross-node communication thereby mitigating the von Neumann bottleneck. Existing work however has generally been limited in that it assumes an accelerator model where kernels are offloaded to SmartNICs with most control tasks left to the CPUs. This leads to frequent waiting reduced performance and scaling challenges. In this work we propose a new distributive data-centric computing framework named FCsN for reconfigurable SmartNIC-based systems. Through a lightweight task circulation execution model and its implementation architecture FCsN allows the complete detaching of NN kernel execution control logic system scheduling and network communication to the SmartNICs. This boosts performance by (i) avoiding control dependency with CPUs and (ii) supporting streaming NN kernel execution and network communication at line rate and in a very fine-grained manner. We demonstrate the efficiency and flexibility of FCsN using various types of neural network kernels and applications including deep neural networks (DNN) and graph neural networks (GNN) as these last are both irregular and data intensive they offer an especially robust demonstration. Evaluations using commonly-used neural network models and graph datasets show that a system with FCsN can achieve 10 × speedups over the MPI-based standard CPU baselines