Implementing a Timing Error-Resilient and Energy-Efficient Near-Threshold Hardware Accelerator for Deep Neural Network Inference

Implementing a Timing Error-Resilient and Energy-Efficient Near-Threshold Hardware Accelerator for Deep Neural Network Inference
复制标题

DOI:
10.3390/jlpea12020032
复制
发表时间:
2022-06
影响因子:
2.1
通讯作者:
N. D. Gundi;Pramesh Pandey;Sanghamitra Roy;Koushik Chakraborty
N. D. Gundi;Pramesh Pandey;Sanghamitra Roy;Koushik Chakraborty
中科院分区:
--
文献类型:
--
作者:
N. D. Gundi;Pramesh Pandey;Sanghamitra Roy;Koushik Chakraborty

文献摘要

相似文献

人工智能(AI)领域日益增长的处理需求导致了深度神经网络(DNN)应用领域特定架构的出现。张量处理单元(TPU)是谷歌的DNN加速器,已经成为领先者,超过了同时代的CPU和GPU,性能提高了15 - 30倍。TPU已部署在Google数据中心,以满足性能需求。然而,TPU的性能增强伴随着猛犸的功耗。为了降低功耗,本文提出了一种工作在近阈值计算(NTC)领域的低功耗TPU--PREDITOR。PREDITOR使用数学分析,通过以特定间隔提升选择性乘法器和累加器单元的电压来减轻无法检测到的时序误差,从而增强NTC TPU的性能,从而确保在低电压下的高推理精度。与领先的误差缓解方案相比,PREDITOR的性能提高了3 - 5倍,精度损失较小。
Increasing processing requirements in the Artificial Intelligence (AI) realm has led to the emergence of domain-specific architectures for Deep Neural Network (DNN) applications. Tensor Processing Unit (TPU), a DNN accelerator by Google, has emerged as a front runner outclassing its contemporaries, CPUs and GPUs, in performance by 15×–30×. TPUs have been deployed in Google data centers to cater to the performance demands. However, a TPU’s performance enhancement is accompanied by a mammoth power consumption. In the pursuit of lowering the energy utilization, this paper proposes PREDITOR—a low-power TPU operating in the Near-Threshold Computing (NTC) realm. PREDITOR uses mathematical analysis to mitigate the undetectable timing errors by boosting the voltage of the selective multiplier-and-accumulator units at specific intervals to enhance the performance of the NTC TPU, thereby ensuring a high inference accuracy at low voltage. PREDITOR offers up to 3×–5× improved performance in comparison to the leading-edge error mitigation schemes with a minor loss in accuracy.