SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration

SSR: Spatial Sequential Hybrid Architecture for Latency Throughput Tradeoff in Transformer Acceleration
复制标题

DOI:
10.1145/3626202.3637569
复制
发表时间:
2024-01
期刊:
Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays
影响因子:
--
通讯作者:
Jinming Zhuang;Zhuoping Yang;Shixin Ji;Heng Huang;Alex K. Jones;Jingtong Hu;Yiyu Shi;Peipei Zhou
Jinming Zhuang;Zhuoping Yang;Shixin Ji;Heng Huang;Alex K. Jones;Jingtong Hu;Yiyu Shi;Peipei Zhou
中科院分区:
其他
文献类型:
--
作者:
Jinming Zhuang;Zhuoping Yang;Shixin Ji;Heng Huang;Alex K. Jones;Jingtong Hu;Yiyu Shi;Peipei Zhou

文献摘要

相似文献

随着芯片计算强度的增加,计算层之间的不匹配和可用的计算资源显着限制了芯片的利用。在此观察结果的驱动下,先前的工作讨论了空间加速器或数据流体系结构,以最大化吞吐量。但是,使用空间加速器可能会增加执行延迟。在这项工作中,我们首先系统地研究了两个执行模型:(1)依次(时间)启动一个单片加速器,(2)(2)空间启动多个加速器。从观察结果来看,我们发现这两个执行模型之间存在延迟的吞吐量权衡,将这两种策略结合在一起可以为我们提供更有效的延迟吞吐量帕累托阵线。为了实现这一目标,我们提出了空间顺序体系结构(SSR)和SSR设计自动化框架,以在部署深度学习推论时一起探索这两种策略。我们使用7NM AMD VERSAL ACAP VCK190板为四个基于端到端变压器的深度学习模型实现SSR加速器。与8nm NVIDIA GPU A10G,16NM AMD FPGAS ZCU102和U250相比,SSR的平均吞吐量增长在不同批次尺寸下的平均吞吐量增长率为2.53倍,35.71倍和14.20倍。平均能效增长分别为8.51倍,6.75倍和21.22倍。与VCK190上的仅一式溶液和仅空间溶液相比,在相同的潜伏期需求和相同的吞吐量要求下,我们的空间序列杂交溶液在相同的潜伏期需求和较低潜伏期下实现了较高的吞吐量。我们还使用SSR分析模型来演示如何使用SSR在其他计算平台上优化解决方案,例如14NM Intel Stratix 10 NX。
With the increase in the computation intensity of the chip, the mismatch between computation layer shapes and the available computation resource significantly limits the utilization of the chip. Driven by this observation, prior works discuss spatial accelerators or dataflow architecture to maximize the throughput. However, using spatial accelerators could potentially increase the execution latency. In this work, we first systematically investigate two execution models: (1) sequentially (temporally) launch one monolithic accelerator, and (2) spatially launch multiple accelerators. From the observations, we find that there is a latency throughput tradeoff between these two execution models, and combining these two strategies together can give us a more efficient latency throughput Pareto front. To achieve this, we propose spatial sequential architecture (SSR) and SSR design automation framework to explore both strategies together when deploying deep learning inference. We use the 7nm AMD Versal ACAP VCK190 board to implement SSR accelerators for four end-to-end transformer-based deep learning models. SSR achieves average throughput gains of 2.53x, 35.71x, and 14.20x under different batch sizes compared to the 8nm Nvidia GPU A10G, 16nm AMD FPGAs ZCU102, and U250. The average energy efficiency gains are 8.51x, 6.75x, and 21.22x, respectively. Compared with the sequential-only solution and spatial-only solution on VCK190, our spatial-sequential-hybrid solutions achieve higher throughput under the same latency requirement and lower latency under the same throughput requirement. We also use SSR analytical models to demonstrate how to use SSR to optimize solutions on other computing platforms, e.g., 14nm Intel Stratix 10 NX.