UNO: Virtualizing and Unifying Nonlinear Operations for Emerging Neural Networks

UNO: Virtualizing and Unifying Nonlinear Operations for Emerging Neural Networks
复制标题

DOI:
10.1109/islped52811.2021.9502473
复制
发表时间:
2021-07
期刊:
2021 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED)
影响因子:
--
通讯作者:
Di Wu;Jingjie Li;Setareh Behroozi;Younghyun Kim
Di Wu;Jingjie Li;Setareh Behroozi;Younghyun Kim
中科院分区:
其他
文献类型:
--
作者:
Di Wu;Jingjie Li;Setareh Behroozi;Younghyun Kim

文献摘要

相似文献

由于线性乘法累加(MAC)运算在传统模型中对能量消耗的贡献占主导地位,因此一直是提高神经网络推理能量效率的主要研究方向。另一方面,除法、乘法和对数等非线性运算在新兴的神经网络模型中正变得越来越重要,但在很大程度上没有得到充分的探索。在本文中,我们提出了UNO,这是一个低面积、低能量的处理元件,它在推理硬件中已有的现成线性MAC单元上虚拟化了非线性运算的泰勒近似。这种虚拟化以统一的、与MAC兼容的方式近似多个非线性运算,以实现动态运行时精度--能量缩放。与基线相比,对于单个操作,我们的方案降低了高达38.4%的能耗,而对于新兴的神经网络模型,在没有可忽略的推理损失的情况下,能量效率提高了高达274.5
Linear multiply-accumulate (MAC) operations have been the main focus of prior efforts in improving the energy efficiency of neural network inference due to their dominant contribution to energy consumption in traditional models. On the other hand, nonlinear operations, such as division, exponentiation, and logarithm, that are becoming increasingly significant in emerging neural network models, have been largely underexplored. In this paper, we propose UNO, a low-area, low-energy processing element that virtualizes the Taylor approximation of nonlinear operations on top of off-the-shelf linear MAC units already present in inference hardware. Such virtualization approximates multiple nonlinear operations in a unified, MAC-compatible manner to achieve dynamic run-time accuracy-energy scaling. Compared to the baseline, our scheme reduces the energy consumption by up to 38.4% for individual operations and increases the energy efficiency by up to 274.5% for emerging neural network models with negligible inference loss