UPTPU: Improving Energy Efficiency of a Tensor Processing Unit through Underutilization Based Power-Gating

UPTPU: Improving Energy Efficiency of a Tensor Processing Unit through Underutilization Based Power-Gating
复制标题

DOI:
10.1109/dac18074.2021.9586224
复制
发表时间:
2021-12
期刊:
2021 58th ACM/IEEE Design Automation Conference (DAC)
影响因子:
--
通讯作者:
Pramesh Pandey;N. D. Gundi;Koushik Chakraborty;Sanghamitra Roy
Pramesh Pandey;N. D. Gundi;Koushik Chakraborty;Sanghamitra Roy
中科院分区:
其他
文献类型:
--
作者:
Pramesh Pandey;N. D. Gundi;Koushik Chakraborty;Sanghamitra Roy

文献摘要

被引文献

相似文献

人工智能的繁荣为神经网络计算带来了过多的特定于领域的架构。谷歌的张量处理单元(TPU),一种深度神经网络(DNN)加速器,已经取代了其数据中心的CPU/GPU,声称推理速度超过15倍。然而,随着人工智能服务的广泛使用,DNN工作负载的前所未有的增长预计基于TPU的数据中心的能耗将不断增加。在这项工作中,我们参数化TPU脉动阵列中的极端硬件利用率不足,并提出UPTPU:一种智能的,低功耗的自适应功率门控范例,为不同的输入批量大小提供惊人的3.5 × - 6.5× TPU的能量效率。
The AI boom is bringing a plethora of domain-specific architectures for Neural Network computations. Google’s Tensor Processing Unit (TPU), a Deep Neural Network (DNN) accelerator, has replaced the CPUs/GPUs in its data centers, claiming more than 15 × rate of inference. However, the unprecedented growth in DNN workloads with the widespread use of AI services projects an increasing energy consumption of TPU based data centers. In this work, we parametrize the extreme hardware underutilization in TPU systolic array and propose UPTPU: an intelligent, dataflow adaptive power-gating paradigm to provide a staggering 3.5 × – 6.5× energy efficiency to TPU for different input batch sizes.