Logarithm-approximate floating-point multiplier is applicable to power-efficient neural network training

Logarithm-approximate floating-point multiplier is applicable to power-efficient neural network training
复制标题

DOI:
10.1016/j.vlsi.2020.05.002
复制
发表时间:
2020-09-01
影响因子:
1.9
通讯作者:
Hashimoto, Masanori
Hashimoto, Masanori
中科院分区:
工程技术4区
文献类型:
--
作者:
Cheng, TaiYu;Masuda, Yukata;Hashimoto, Masanori

文献摘要

被引文献

相似文献

最近,新兴的“边缘计算”将数据和服务从云端移动到附近的边缘服务器,以实现短延迟和宽带宽,并解决隐私问题。然而,边缘服务器通常嵌入GPU处理器,由于功率和尺寸的限制,高度要求节能神经网络(NN)训练解决方案。此外,根据在NN训练中计算的梯度值的宽动态范围的性质,浮点表示更适合。本文提出在神经网络(NN)训练引擎中采用对数近似乘法器(LAM)进行乘法累加(MAC)计算,其中LAM将浮点乘法近似为定点加法,从而具有更小的延迟、更少的门和更低的功耗。我们在两个平台上展示了LAM的效率,这两个平台是专用的NN训练硬件和开源GPU设计。与应用精确乘法器的NN训练相比,我们为二维分类数据集实现的NN训练引擎在功率和面积上分别实现了10%的速度提升和2.3倍的效率提升。LAM还与传统的位宽缩放(BWS)高度兼容。当在五个测试数据集中应用BWS与LAM时,所实现的训练引擎实现了超过4.9倍的功率效率改进,最多1%的准确性下降,其中2.2倍的改进来自LAM。此外,LAM的优点可以在处理器中利用。嵌入LAM的GPU设计执行NN训练工作负载,在FPGA中实现,呈现1.32倍的功率效率提高,而LAM + BWS的提高达到1.54倍。最后,评估了基于LAM的深层NN训练。多达4个隐藏层NN,基于LAM的训练实现了与精确乘法器高度相当的准确性,即使具有积极的BWS。
Recently, emerging "edge computing" moves data and services from the cloud to nearby edge servers to achieve short latency and wide bandwidth, and solve privacy concerns. However, edge servers, often embedded with GPU processors, highly demand a solution for power-efficient neural network (NN) training due to the limitation of power and size. Besides, according to the nature of the broad dynamic range of gradient values computed in NN training, floating-point representation is more suitable. This paper proposes to adopt a logarithm-approximate multiplier (LAM) for multiply-accumulate (MAC) computation in neural network (NN) training engines, where LAM approximates a floating-point multiplication as a fixed-point addition, resulting in smaller delay, fewer gates, and lower power consumption. We demonstrate the efficiency of LAM in two platforms, which are dedicated NN training hardware, and open-source GPU design. Compared to the NN training applying the exact multiplier, our implementation of the NN training engine for a 2-D classification dataset achieves 10% speed-up and 2.3X efficiency improvement in power and area, respectively. LAM is also highly compatible with conventional bit-width scaling (BWS). When BWS is applied with LAM in five test datasets, the implemented training engines achieve more than 4.9X power efficiency improvement, with at most 1% accuracy degradation, where 2.2X improvement originates from LAM. Also, the advantage of LAM can be exploited in processors. A GPU design embedded with LAM executing an NN-training workload, which is implemented in an FPGA, presents 1.32X power efficiency improvement, and the improvement reaches 1.54X with LAM + BWS. Finally, LAM-based training in deeper NN is evaluated. Up to 4-hidden layer NN, LAM-based training achieves highly comparable accuracy as that of the accurate multiplier, even with aggressive BWS.