E.T.: Re-Thinking Self-Attention for Transformer Models on GPUs

E.T.: Re-Thinking Self-Attention for Transformer Models on GPUs
复制标题

DOI:
10.1145/3458817.3476138
复制
发表时间:
2021-11
期刊:
SC21: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Shiyang Chen;Shaoyi Huang;Santosh Pandey;Bingbing Li;G. Gao;Long Zheng;Caiwen Ding;Hang Liu
Shiyang Chen;Shaoyi Huang;Santosh Pandey;Bingbing Li;G. Gao;Long Zheng;Caiwen Ding;Hang Liu
中科院分区:
其他
文献类型:
--
作者:
Shiyang Chen;Shaoyi Huang;Santosh Pandey;Bingbing Li;G. Gao;Long Zheng;Caiwen Ding;Hang Liu

文献摘要

相似文献

基于变压器的深度学习模型已成为一种无处不在的工具,以推动各种自然语言处理(NLP)超出其准确性上限的相关任务。但是,这些模型还遭受了两个明显的挑战,即巨大的模型大小和延长的周转时间。为此,我们介绍了E.T.这重新考虑了GPU上的变压器模型的自我注意计算,其贡献是:首先,我们引入了一种新型的自我发项式结构,该结构包括两个量身定制的自我发项式操作员,具有相应的序列长度意识到的优化,并重新排序了优化。其次,我们提出了一种注意力感知的修剪设计,该设计明智地使用各种修剪算法来减少更多计算,因此实现了更短的周转时间。对于修剪算法,我们不仅修改了现有的修剪算法,还为变压器模型量身定制了新的算法。综上所述,我们评估了E.T.在变压器,Bertbase和Distilbert的各种基准中,其中E.T.提出优于主流项目的卓越性能,包括流行的NVIDIA Enterprise解决方案,即Tensorrt和ForterTransFormer。
Transformer-based deep learning models have become a ubiquitous vehicle to drive a variety of Natural Language Processing (NLP) related tasks beyond their accuracy ceiling. However, these models also suffer from two pronounced challenges, that is, gigantic model size and prolonged turnaround time. To this end, we introduce E.T. that rE-thinks self-attention computation for Transformer models on GPUs with the following contributions: First, we introduce a novel self-attention architecture, which encompasses two tailored self-attention operators with corresponding sequence length-aware optimizations, and operation reordering optimizations. Second, we present an attention-aware pruning design which judiciously uses various pruning algorithms to reduce more computations hence achieves significantly shorter turnaround time. For the pruning algorithms, we not only revamp the existing pruning algorithms, but also tailor new ones for transformer models. Taken together, we evaluate E.T. across a variety of benchmarks for Transformer, BERTBASE and DistilBERT, where E.T. presents superior performance over the mainstream projects, including the popular Nvidia Enterprise solutions, i.e., TensorRT and FasterTransformer.