Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design

Adaptable Butterfly Accelerator for Attention-based NNs via Hardware and Algorithm Co-design
复制标题

DOI:
10.1109/micro56248.2022.00050
复制
发表时间:
2022-09
期刊:
2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Hongxiang Fan;Thomas C. P. Chau;Stylianos I. Venieris;Royson Lee;Alexandros Kouris;W. Luk;N. Lane;M. Abdelfattah
Hongxiang Fan;Thomas C. P. Chau;Stylianos I. Venieris;Royson Lee;Alexandros Kouris;W. Luk;N. Lane;M. Abdelfattah
中科院分区:
其他
文献类型:
--
作者:
Hongxiang Fan;Thomas C. P. Chau;Stylianos I. Venieris;Royson Lee;Alexandros Kouris;W. Luk;N. Lane;M. Abdelfattah

文献摘要

被引文献

相似文献

基于注意力的神经网络已经在许多人工智能任务中变得普遍。尽管它们具有出色的算法性能,但注意力机制和前馈网络(FFN)的使用需要过多的计算和存储器资源,这通常会损害它们的硬件性能。虽然已经引入了各种稀疏变体,但大多数方法仅关注于减轻算法级别上的注意力的二次缩放,而没有明确考虑将其方法映射到真实的硬件设计上的效率。此外,大多数努力只集中在注意力机制或FFN,但没有联合优化这两个部分,导致大多数当前的设计缺乏可扩展性时,处理不同的输入长度。本文从硬件的角度系统地考虑了不同变体中的稀疏模式。在算法层面上,我们提出了FABNet,一个硬件友好的变体,采用统一的蝴蝶稀疏模式来近似注意力机制和FFN。在硬件层面上,提出了一种新的自适应蝶形加速器,可以在运行时通过专用的硬件控制配置,以加速不同的蝶形层使用一个统一的硬件引擎。在Long-Range-竞技场数据集上,FABNet实现了与普通Transformer相同的精度,同时减少了10$\sim66\times$的计算量和2$\sim22\times$的参数数量。通过联合优化算法和硬件,我们的基于FPGA的蝶形加速器实现了14.2$\sim23.2\times$的加速比,超过了标准化为相同计算预算的最先进的加速器。与Raspberry Pi 4和Jetson Nano上优化的CPU和GPU设计相比,在相同的功耗预算下,我们的系统速度分别提高了273.8\times $和15.1\times $
Attention-based neural networks have become pervasive in many AI tasks. Despite their excellent algorithmic performance, the use of the attention mechanism and feedforward network (FFN) demands excessive computational and memory resources, which often compromises their hardware performance. Although various sparse variants have been introduced, most approaches only focus on mitigating the quadratic scaling of attention on the algorithm level, without explicitly considering the efficiency of mapping their methods on real hardware designs. Furthermore, most efforts only focus on either the attention mechanism or the FFNs but without jointly optimizing both parts, causing most of the current designs to lack scalability when dealing with different input lengths. This paper systematically considers the sparsity patterns in different variants from a hardware perspective. On the algorithmic level, we propose FABNet, a hardware-friendly variant that adopts a unified butterfly sparsity pattern to approximate both the attention mechanism and the FFNs. On the hardware level, a novel adaptable butterfly accelerator is proposed that can be configured at runtime via dedicated hardware control to accelerate different butterfly layers using a single unified hardware engine. On the Long-Range-Arena dataset, FABNet achieves the same accuracy as the vanilla Transformer while reducing the amount of computation by 10$\sim66\times$ and the number of parameters 2$\sim22\times$. By jointly optimizing the algorithm and hardware, our FPGA-based butterfly accelerator achieves 14.2$\sim23.2\times$ speedup over state-of-the-art accelerators normalized to the same computational budget. Compared with optimized CPU and GPU designs on Raspberry Pi 4 and Jetson Nano, our system is up to $273.8\times$ and $15.1\times$ faster under the same power budget