Output-Directed Dynamic Quantization for DNN Acceleration
Output-Directed Dynamic Quantization for DNN Acceleration
复制标题
DOI:
10.1145/3605573.3605580
复制
发表时间:
2023-08
期刊:
影响因子:
--
通讯作者:
Beilei Jiang;Xianwei Cheng;Yuan Li;Jocelyn Zhang;Song Fu;Qing Yang;Mingxiong Liu;Alejandro Olvera
中科院分区:
文献类型:
--
作者:
Beilei Jiang;Xianwei Cheng;Yuan Li;Jocelyn Zhang;Song Fu;Qing Yang;Mingxiong Liu;Alejandro Olvera
Quantization is an effective technique for reducing the number of computations and improving the performance of deep neural networks (DNNs). Weight quantization is popular because weights can be trained beforehand. However, weight quantization only targets the kernel weights and ignores the sensitivity of input features, which can lead to reduced accuracy. Fine-grained input quantization has gained attention as a way to speed up DNNs while maintaining accuracy. Existing approaches determine computation precision based on input sensitivity but do not effectively reduce computations for insensitive outputs or retain the precision of sensitive outputs. These limitations motivate us to develop an output-directed dynamic quantization method named ODQ in this paper. ODQ is a two-stage DNN quantization scheme designed to improve performance, reduce energy consumption, and maintain and often improve accuracy, compared with existing quantization methods. Specifically, inputs and weights go through sensitivity prediction and result generation. The high-order 2 bits of input and weight are used to predict output sensitivity. Result generation is performed only for predicted sensitive outputs. We designed an FPGA accelerator to optimize ODQ quantization performance for DNNs. We implement a prototype of ODQ and evaluate its performance using several state-of-the-art DNNs. Compared with a state-of-the-art input-directed quantization approach, ODQ achieves a 67.6% performance speedup and a 66.9% energy saving, with minimal accuracy degradation (≤ 0.6%).