Accelerate Scientific Deep Learning Models on Heterogeneous Computing Platform with FPGA

Accelerate Scientific Deep Learning Models on Heterogeneous Computing Platform with FPGA
复制标题

利用 FPGA 在异构计算平台上加速科学深度学习模型

DOI:
10.1051/epjconf/202024509014
复制
发表时间:
2020
影响因子:
--
通讯作者:
C. Jiang, D. Ojika
C. Jiang, D. Ojika
中科院分区:
--
文献类型:
--
作者:
C. Jiang, D. Ojika

文献摘要

相似文献

人工智能和深度学习几乎在涉及大数据分析的每个领域都经历了爆炸式增长。使用深度神经网络(DNN)的深度学习在此类科学数据分析应用中显示出巨大的前景。然而,传统的基于CPU的顺序计算,没有特殊的指令,已经不能满足任务关键型应用程序的要求,这是计算密集型,需要低延迟和高吞吐量。异构计算(HGC)的CPU与GPU、FPGA和其他科学目标加速器集成,提供了加速DNN的独特功能。佛罗里达大学的SHREC 1、CERN Openlab、劳伦斯伯克利国家实验室的NERSC 2、Dell EMC和英特尔的合作研究人员正在研究异构计算(HGC)在使用DNN模型解决科学问题方面的应用。本文重点介绍了使用FPGA来加速HGC工作流程的推理阶段。我们目前的案例研究和结果推断国家的最先进的DNN模型的科学数据分析,使用英特尔发行的OpenVINO,运行在英特尔可编程加速卡(PAC)配备了阿里亚10 GX FPGA。使用英特尔深度学习加速(DLA)开发套件优化现有的FPGA原语并开发新的原语,我们能够加速正在研究的科学DNN模型,单个Arria 10 FPGA相对于服务器级Skylake CPU的单核(单线程)的加速比从2.46倍提高到9.59倍。
AI and deep learning are experiencing explosive growth in almost every domain involving analysis of big data. Deep learning using Deep Neural Networks (DNNs) has shown great promise for such scientific data analysis applications. However, traditional CPU-based sequential computing without special instructions can no longer meet the requirements of mission-critical applications, which are compute-intensive and require low latency and high throughput. Heterogeneous computing (HGC), with CPUs integrated with GPUs, FPGAs, and other science-targeted accelerators, offers unique capabilities to accelerate DNNs. Collaborating researchers at SHREC1at the University of Florida, CERN Openlab, NERSC2at Lawrence Berkeley National Lab, Dell EMC, and Intel are studying the application of heterogeneous computing (HGC) to scientific problems using DNN models. This paper focuses on the use of FPGAs to accelerate the inferencing stage of the HGC workflow. We present case studies and results in inferencing state-of-the-art DNN models for scientific data analysis, using Intel distribution of OpenVINO, running on an Intel Programmable Acceleration Card (PAC) equipped with an Arria 10 GX FPGA. Using the Intel Deep Learning Acceleration (DLA) development suite to optimize existing FPGA primitives and develop new ones, we were able accelerate the scientific DNN models under study with a speedup from 2.46x to 9.59x for a single Arria 10 FPGA against a single core (single thread) of a server-class Skylake CPU.