Output-Directed Dynamic Quantization for DNN Acceleration

Output-Directed Dynamic Quantization for DNN Acceleration
复制标题

DOI:
10.1145/3605573.3605580
复制
发表时间:
2023-08
期刊:
Proceedings of the 52nd International Conference on Parallel Processing
影响因子:
--
通讯作者:
Beilei Jiang;Xianwei Cheng;Yuan Li;Jocelyn Zhang;Song Fu;Qing Yang;Mingxiong Liu;Alejandro Olvera
Beilei Jiang;Xianwei Cheng;Yuan Li;Jocelyn Zhang;Song Fu;Qing Yang;Mingxiong Liu;Alejandro Olvera
中科院分区:
其他
文献类型:
--
作者:
Beilei Jiang;Xianwei Cheng;Yuan Li;Jocelyn Zhang;Song Fu;Qing Yang;Mingxiong Liu;Alejandro Olvera

文献摘要

相似文献

量化是减少计算量并提高深度神经网络(DNN)性能的有效技术。权重量化很受欢迎,因为可以预先训练权重。然而,权重量化仅针对核权重,忽略了输入特征的敏感性,这可能导致准确性降低。细粒度输入量化作为一种​​在保持准确性的同时加速 DNN 的方法而受到关注。现有方法根据输入灵敏度确定计算精度,但不能有效减少对不敏感输出的计算或保留敏感输出的精度。这些限制促使我们在本文中开发一种名为 ODQ 的输出导向动态量化方法。 ODQ 是一种两阶段 DNN 量化方案,旨在与现有量化方法相比提高性能、降低能耗、保持并经常提高精度。具体来说,输入和权重经过敏感性预测和结果生成。输入和权重的高位 2 位用于预测输出灵敏度。仅针对预测的敏感输出执行结果生成。我们设计了一款 FPGA 加速器来优化 DNN 的 ODQ 量化性能。我们实现了 ODQ 原型,并使用几种最先进的 DNN 评估其性能。与最先进的输入定向量化方法相比,ODQ 实现了 67.6% 的性能加速和 66.9% 的节能,并且精度下降最小(≤ 0.6%)。
Quantization is an effective technique for reducing the number of computations and improving the performance of deep neural networks (DNNs). Weight quantization is popular because weights can be trained beforehand. However, weight quantization only targets the kernel weights and ignores the sensitivity of input features, which can lead to reduced accuracy. Fine-grained input quantization has gained attention as a way to speed up DNNs while maintaining accuracy. Existing approaches determine computation precision based on input sensitivity but do not effectively reduce computations for insensitive outputs or retain the precision of sensitive outputs. These limitations motivate us to develop an output-directed dynamic quantization method named ODQ in this paper. ODQ is a two-stage DNN quantization scheme designed to improve performance, reduce energy consumption, and maintain and often improve accuracy, compared with existing quantization methods. Specifically, inputs and weights go through sensitivity prediction and result generation. The high-order 2 bits of input and weight are used to predict output sensitivity. Result generation is performed only for predicted sensitive outputs. We designed an FPGA accelerator to optimize ODQ quantization performance for DNNs. We implement a prototype of ODQ and evaluate its performance using several state-of-the-art DNNs. Compared with a state-of-the-art input-directed quantization approach, ODQ achieves a 67.6% performance speedup and a 66.9% energy saving, with minimal accuracy degradation (≤ 0.6%).