EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference

EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
复制标题

DOI:
10.1145/3466752.3480095
复制
发表时间:
2020-11
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Thierry Tambe;Coleman Hooper;Lillian Pentecost;Tianyu Jia;En-Yu Yang;M. Donato;Victor Sanh;P. Whatmough;Alexander M. Rush;D. Brooks;Gu-Yeon Wei
Thierry Tambe;Coleman Hooper;Lillian Pentecost;Tianyu Jia;En-Yu Yang;M. Donato;Victor Sanh;P. Whatmough;Alexander M. Rush;D. Brooks;Gu-Yeon Wei
中科院分区:
其他
文献类型:
--
作者:
Thierry Tambe;Coleman Hooper;Lillian Pentecost;Tianyu Jia;En-Yu Yang;M. Donato;Victor Sanh;P. Whatmough;Alexander M. Rush;D. Brooks;Gu-Yeon Wei

文献摘要

被引文献

相似文献

基于变压器的语言模型(例如BERT)为多种自然语言处理(NLP)任务提供了显着的准确性,但他们的计算和内存要求使它们挑战到具有严格的延迟需求的资源受限的边缘平台Edgebert,用于多任务NLP的延迟感知能量优化的深入算法 - 硬件提前出口预测以执行动态电压缩放(DVF),以句子粒度进行最小的能量消耗,同时通过使用适应性注意跨度的校准组合来进一步缓解规定的目标延迟。选择性网络修剪和浮点量化。计算设置,我们专门使用12nm可扩展的硬件加速器系统,集成了快速开关的低丢弃电压调节器(LDO),一个全数字的相位锁定环(ADPLL)以及高密度嵌入式的非volted嵌入式的非volted嵌入式记忆(envms)其中仔细存储了共享多任务参数的稀疏浮点位编码。与传统推断相比,最多可产生7×,2.5×和53倍的能量,而没有早期停止,延迟无遇到的早期退出方法以及对NVIDIA JETSON TEGRA X2移动GPU的CUDA适应。
Transformer-based language models such as BERT provide significant accuracy improvement to a multitude of natural language processing (NLP) tasks. However, their hefty computational and memory demands make them challenging to deploy to resource-constrained edge platforms with strict latency requirements. We present EdgeBERT, an in-depth algorithm-hardware co-design for latency-aware energy optimizations for multi-task NLP. EdgeBERT employs entropy-based early exit predication in order to perform dynamic voltage-frequency scaling (DVFS), at a sentence granularity, for minimal energy consumption while adhering to a prescribed target latency. Computation and memory footprint overheads are further alleviated by employing a calibrated combination of adaptive attention span, selective network pruning, and floating-point quantization. Furthermore, in order to maximize the synergistic benefits of these algorithms in always-on and intermediate edge computing settings, we specialize a 12nm scalable hardware accelerator system, integrating a fast-switching low-dropout voltage regulator (LDO), an all-digital phase-locked loop (ADPLL), as well as, high-density embedded non-volatile memories (eNVMs) wherein the sparse floating-point bit encodings of the shared multi-task parameters are carefully stored. Altogether, latency-aware multi-task NLP inference acceleration on the EdgeBERT hardware system generates up to 7 ×, 2.5 ×, and 53 × lower energy compared to the conventional inference without early stopping, the latency-unbounded early exit approach, and CUDA adaptations on an Nvidia Jetson Tegra X2 mobile GPU, respectively.