From Algorithm to Module: Adaptive and Energy-Efficient Quantization Method for Edge Artificial Intelligence in IoT Society

From Algorithm to Module: Adaptive and Energy-Efficient Quantization Method for Edge Artificial Intelligence in IoT Society
复制标题

DOI:
10.1109/tii.2022.3223222
复制
发表时间:
2023-08
影响因子:
12.3
通讯作者:
Tao Li;Yitao Ma;T. Endoh
Tao Li;Yitao Ma;T. Endoh
中科院分区:
计算机科学1区
文献类型:
--
作者:
Tao Li;Yitao Ma;T. Endoh

文献摘要

相似文献

下一代工业边缘人工智能(AI)应用无疑将出现在包含各种传感器、处理器和功能模块的高能效、高度集成的平台上。将数据量化模块嵌入边缘AI芯片,连接传感器、处理器和功能模块,对于实现各种数据表示格式的自适应转换至关重要。本文提出了一种新的自适应低功耗量化技术,并系统验证了其从算法到硬件模块的有效性,用于工业物联网应用,涵盖自动驾驶汽车的精确导航和利用深度神经网络(dnn)的准确分类。提出的量化方法融合了自适应浮点到定点二进制数的转换函数和自适应基点确定函数,保证了定点输入到边缘人工智能模块的足够分辨率和最小误差损失。实验结果表明,量化误差对捷联惯导系统的导航解和dnn的top-1和top-5分类精度(分别为10$^{-8}$和10$^{-7}$)的影响可以忽略。此外,根据所提出的量化技术,设计、合成和路由了量子化乘法器(QoM)硬件模块。仿真结果表明,QoM的功耗和面积分别为0.1 mW和649.552 $\mu \text{m}^{2}$,分别占其总功耗和总面积的5%和14.15%。采用我们提出的片上量化技术,量化dnn参数所需的时间比现有的基准片外量化方法缩短了1142倍。
Next-generation industrial edge artificial intelligence (AI) applications will undoubtedly emerge on energy-efficient, highly-integrated platforms incorporating various sensors, processors, and functional modules. Embedding the data quantization module onto edge AI chips, connecting sensors, processors, and functional modules is critical to achieving adaptive transformation of diverse data representation formats. This paper proposes a novel adaptive and low-power quantization technique and systematically validates its effectiveness from algorithm to hardware module for industrial IoT applications, covering precise navigation for autonomous vehicles and accurate classification utilizing deep neural networks (DNNs). The proposed quantization method merges an adaptive conversion function from floating-point to fixed-point binaries and an adaptive radix-point determination function, ensuring adequate resolution and minimal error loss of the fixed-point inputs to the edge AI modules. The experimental results demonstrate that the quantization error in the proposed quantization technique contributes negligible errors to the navigation solutions of the strapdown inertial navigation system and the DNNs' top-1 and top-5 classification accuracy (on the order of 10$^{-8}$ and 10$^{-7}$). Moreover, a quantization-on-multiplier (QoM) hardware module is designed, synthesized, and routed in accordance with the proposed quantization technique. The simulation results indicate that the QoM's power consumption and area are 0.1 mW and 649.552 $\mu \text{m}^{2}$, accounting for 5% and 14.15% of its total energy consumption and area, respectively. With our proposed on-chip quantization technique, the time required to quantize the parameters of DNNs is up to 1142 times shorter than with the existing benchmark off-chip quantization approaches.