CHARM: Composing Heterogeneous AcceleRators for Matrix Multiply on Versal ACAP Architecture

CHARM: Composing Heterogeneous AcceleRators for Matrix Multiply on Versal ACAP Architecture
复制标题

CHARM:在 Versal ACAP 架构上组合用于矩阵乘法的异构加速器

DOI:
10.1145/3543622.3573210
复制
发表时间:
2023
期刊:
Proceedings of the 2023 ACM/SIGDA International Symposium on Field Programmable Gate Arrays
影响因子:
--
通讯作者:
Hu, Jingtong
Hu, Jingtong
中科院分区:
--
文献类型:
--
作者:
Zhuang, Jinming;Lau, Jason;Ye, Hanchen;Yang, Zhuoping;Du, Yubo;Lo, Jack;Denolf, Kristof;Neuendorffer, Stephen;Jones, Alex;Hu, Jingtong

文献摘要

参考文献

被引文献

相似文献

密集矩阵乘法(MM)是深度学习应用中使用最多的内核之一。为了科普这些应用的高计算需求,以FPGA和专用ASIC加速器为特色的异构架构已成为有前途的平台。例如,AMD/Xilinx Versal ACAP架构将通用CPU内核和可编程逻辑(PL)与针对AI/ML优化的AI引擎处理器(AIE)相结合。在1 GHz下执行的400个AI Engine处理器阵列理论上可以为32位浮点(fp 32)数据提供高达6.4 TFLOP的性能。然而,机器学习模型通常包含大型和小型MM操作。虽然大型MM操作可以跨多个核高效并行化,但小型MM操作通常不能。在我们的调查中,我们观察到在Versal ACAP中的大型单片MM加速器上执行BERT自然语言处理模型中的一些小MM层,实现了不到理论峰值性能的5%。因此,一个关键问题出现了:我们如何设计加速器,以充分利用有限的通信带宽下的丰富的计算资源,为端到端的应用程序与不同大小的多个MM层?我们确定了最大的系统吞吐量的瓶颈造成的不匹配的大规模计算资源的一个单片加速器和各种MM层的小尺寸的应用程序。为了解决这个问题,我们提出了CHARM框架来组合多个不同的MM加速器架构,同时在一个应用程序中的不同层工作。CHARM包括分析模型,指导设计空间探索,以确定加速器分区和层调度。为了方便系统设计,CHARM自动生成代码,实现彻底的板载设计验证。我们在AMD/Xilinx Versal ACAP VCK 190评估板上为四种不同的深度学习应用部署了CHARM框架,包括BERT、ViT、NCF和MLP。实验结果表明,在BERT、ViT、NCF和MLP四种情况下,我们分别获得了1.46、1.61、1.74和2.94 TFLOPs的推理吞吐量,与单块加速器相比,分别获得了5.40倍、32.51倍、1.00倍和1.00倍的吞吐量增益。
Dense matrix multiply (MM) serves as one of the most heavily used kernels in deep learning applications. To cope with the high computation demands of these applications, heterogeneous architectures featuring both FPGA and dedicated ASIC accelerators have emerged as promising platforms. For example, the AMD/Xilinx Versal ACAP architecture combines general-purpose CPU cores and programmable logic (PL) with AI Engine processors (AIE) optimized for AI/ML. An array of 400 AI Engine processors executing at 1 GHz can theoretically provide up to 6.4 TFLOPs performance for 32-bit floating-point (fp32) data. However, machine learning models often contain both large and small MM operations. While large MM operations can be parallelized efficiently across many cores, small MM operations typically cannot. In our investigation, we observe that executing some small MM layers from the BERT natural language processing model on a large, monolithic MM accelerator in Versal ACAP achieved less than 5% of the theoretical peak performance. Therefore, one key question arises: How can we design accelerators to fully use the abundant computation resources under limited communication bandwidth for end-to-end applications with multiple MM layers of diverse sizes?We identify the biggest system throughput bottleneck resulting from the mismatch of massive computation resources of one monolithic accelerator and the various MM layers of small sizes in the application. To resolve this problem, we propose the CHARM framework to composemultiple diverse MM accelerator architecturesworking concurrently towards different layers within one application. CHARM includes analytical models which guide design space exploration to determine accelerator partitions and layer scheduling. To facilitate the system designs, CHARM automatically generates code, enabling thorough onboard design verification. We deploy the CHARM framework for four different deep learning applications, including BERT, ViT, NCF, MLP, on the AMD/Xilinx Versal ACAP VCK190 evaluation board. Our experiments show that we achieve 1.46 TFLOPs, 1.61 TFLOPs, 1.74 TFLOPs, and 2.94 TFLOPs inference throughput for BERT, ViT, NCF, MLP, respectively, which obtain 5.40x, 32.51x, 1.00x and 1.00x throughput gains compared to one monolithic accelerator.
全流水线的能源效率:矩阵乘法的案例研究
DOI: --
发表时间: 2016
期刊: IEEE Symposium on Field-Programmable Custom Computing Machines
影响因子: --
作者:
Peipei Zhou;Hyunseok Park;Zhenman Fang;J. Cong;A. DeHon
通讯作者: A. DeHon
DOI: --
发表时间: 2012
期刊: 2012 SC Companion: High Performance Computing, Networking Storage and Analysis
影响因子: --
作者:
J. Demmel
通讯作者: J. Demmel
DOI: 10.1145/3431920.3439292
发表时间: 2021-02
期刊: The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
影响因子: --
作者:
Jie Wang;Licheng Guo;J. Cong
通讯作者: Jie Wang;Licheng Guo;J. Cong
CROW:一种用于提高 DRAM 性能、能源效率和可靠性的低成本基板
DOI: 10.1145/3307650.3322231
发表时间: 2019
期刊: 2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者:
Hasan Hassan;Minesh Patel;Jeremie S. Kim;A. G. Yaglikçi;Nandita Vijaykumar;Nika Mansouri;Saugata Ghose;O. Mutlu
通讯作者: O. Mutlu