CHARM: Composing Heterogeneous AcceleRators for Matrix Multiply on Versal ACAP Architecture
CHARM: Composing Heterogeneous AcceleRators for Matrix Multiply on Versal ACAP Architecture
复制标题
CHARM:在 Versal ACAP 架构上组合用于矩阵乘法的异构加速器
DOI:
10.1145/3543622.3573210
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Hu, Jingtong
中科院分区:
文献类型:
--
作者:
Zhuang, Jinming;Lau, Jason;Ye, Hanchen;Yang, Zhuoping;Du, Yubo;Lo, Jack;Denolf, Kristof;Neuendorffer, Stephen;Jones, Alex;Hu, Jingtong
Dense matrix multiply (MM) serves as one of the most heavily used kernels in deep learning applications. To cope with the high computation demands of these applications, heterogeneous architectures featuring both FPGA and dedicated ASIC accelerators have emerged as promising platforms. For example, the AMD/Xilinx Versal ACAP architecture combines general-purpose CPU cores and programmable logic (PL) with AI Engine processors (AIE) optimized for AI/ML. An array of 400 AI Engine processors executing at 1 GHz can theoretically provide up to 6.4 TFLOPs performance for 32-bit floating-point (fp32) data. However, machine learning models often contain both large and small MM operations. While large MM operations can be parallelized efficiently across many cores, small MM operations typically cannot. In our investigation, we observe that executing some small MM layers from the BERT natural language processing model on a large, monolithic MM accelerator in Versal ACAP achieved less than 5% of the theoretical peak performance. Therefore, one key question arises: How can we design accelerators to fully use the abundant computation resources under limited communication bandwidth for end-to-end applications with multiple MM layers of diverse sizes?We identify the biggest system throughput bottleneck resulting from the mismatch of massive computation resources of one monolithic accelerator and the various MM layers of small sizes in the application. To resolve this problem, we propose the CHARM framework to composemultiple diverse MM accelerator architecturesworking concurrently towards different layers within one application. CHARM includes analytical models which guide design space exploration to determine accelerator partitions and layer scheduling. To facilitate the system designs, CHARM automatically generates code, enabling thorough onboard design verification. We deploy the CHARM framework for four different deep learning applications, including BERT, ViT, NCF, MLP, on the AMD/Xilinx Versal ACAP VCK190 evaluation board. Our experiments show that we achieve 1.46 TFLOPs, 1.61 TFLOPs, 1.74 TFLOPs, and 2.94 TFLOPs inference throughput for BERT, ViT, NCF, MLP, respectively, which obtain 5.40x, 32.51x, 1.00x and 1.00x throughput gains compared to one monolithic accelerator.
登录
查看更多内容
DOI:
--
发表时间:
2016
期刊:
IEEE Symposium on Field-Programmable Custom Computing Machines
影响因子:
--
作者:
Peipei Zhou;Hyunseok Park;Zhenman Fang;J. Cong;A. DeHon
通讯作者:
A. DeHon
DOI:
--
发表时间:
2012
期刊:
2012 SC Companion: High Performance Computing, Networking Storage and Analysis
影响因子:
--
作者:
J. Demmel
通讯作者:
J. Demmel
DOI:
10.1145/3431920.3439292
发表时间:
2021-02
期刊:
The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
影响因子:
--
作者:
Jie Wang;Licheng Guo;J. Cong
通讯作者:
Jie Wang;Licheng Guo;J. Cong
DOI:
10.1145/3307650.3322231
发表时间:
2019
期刊:
2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
作者:
Hasan Hassan;Minesh Patel;Jeremie S. Kim;A. G. Yaglikçi;Nandita Vijaykumar;Nika Mansouri;Saugata Ghose;O. Mutlu
通讯作者:
O. Mutlu