A Survey on Sparsity Exploration in Transformer-Based Accelerators

A Survey on Sparsity Exploration in Transformer-Based Accelerators
复制标题

DOI:
10.3390/electronics12102299
复制
发表时间:
2023-05-19
期刊:
影响因子:
2.9
通讯作者:
Chen, Lizhong
Chen, Lizhong
中科院分区:
工程技术3区
文献类型:
--
作者:
Fuad, Kazi Ahmed Asif;Chen, Lizhong

文献摘要

相似文献

由于能够处理更长的标记序列并更高效地支持并行处理,Transformer模型已成为众多自然语言处理和计算机视觉应用中的前沿技术。然而,Transformer模型的训练和推理计算成本高昂且内存需求大。与此同时,利用深度学习模型中的稀疏性已被证明是缓解计算难题以及助力将大型模型适配到边缘设备的有效方法。鉴于高性能中央处理器(CPU)和图形处理器(GPU)通常在探索底层稀疏性方面灵活性不足,人们已针对Transformer模型提出了许多专用硬件加速器。本文全面综述了为探索稀疏性以实现计算和内存优化而提出的Transformer硬件加速器。我们根据利用稀疏性的策略对现有研究进行分类,并指出这些策略的优缺点。基于分析结果,我们为未来改进Transformer硬件加速器有效稀疏执行的研究指明了有前景的方向并给出建议。
Transformer models have emerged as the state-of-the-art in many natural language processing and computer vision applications due to their capability of attending to longer sequences of tokens and supporting parallel processing more efficiently. Nevertheless, the training and inference of transformer models are computationally expensive and memory intensive. Meanwhile, utilizing the sparsity in deep learning models has proven to be an effective approach to alleviate the computation challenge as well as help to fit large models in edge devices. As high-performance CPUs and GPUs are generally not flexible enough to explore low-level sparsity, a number of specialized hardware accelerators have been proposed for transformer models. This paper provides a comprehensive review of hardware transformer accelerators that have been proposed to explore sparsity for computation and memory optimizations. We classify existing works based on the strategies of utilizing sparsity and identify their pros and cons in those strategies. Based on our analysis, we point out promising directions and recommendations for future works on improving the effective sparse execution of transformer hardware accelerators.