ReLeQ : A Reinforcement Learning Approach for Automatic Deep Quantization of Neural Networks.

ReLeQ : A Reinforcement Learning Approach for Automatic Deep Quantization of Neural Networks.
复制标题

DOI:
10.1109/mm.2020.3009475
复制
发表时间:
2020-09
期刊:
影响因子:
3.6
通讯作者:
Yazdanbakhsh A
Yazdanbakhsh A
中科院分区:
计算机科学3区
文献类型:
--
作者:
Elthakeb AT;Pilligundla P;Mireshghallah F;Esmaeilzadeh H;Yazdanbakhsh A

文献摘要

被引文献

相似文献

深度量化(八位以下)可以通过减少网络编码的位宽来显着减少 DNN 计算和存储。然而,如果没有艰苦的手动工作,这种深度量化可能会导致严重的精度损失,使其实用性受到质疑。我们提出了一种系统方法来解决这个问题,通过端到端深度强化学习框架(ReLeQ)自动化发现位宽的过程。该框架利用近端策略优化的样本效率来探索层的位宽可能分配的指数级大空间。我们展示了 ReLeQ 如何平衡速度和质量,并为各种深度网络的量化提供异构位宽分配,以最小的精度损失(≤ 0.3% 损失),同时最大限度地减少计算和存储成本。借助这些 DNN,ReLeQ 使传统硬件和定制 DNN 加速器能够比 8 位执行实现 2.2 倍的加速。
Deep Quantization (below eight bits) can significantly reduce the DNN computation and storage by decreasing the bitwidth of network encodings. However, without arduous manual effort, this deep quantization can lead to significant accuracy loss, leaving it in a position of questionable utility. We propose a systematic approach to tackle this problem, by automating the process of discovering the bitwidths through an end-to-end deep reinforcement learning framework (ReLeQ). This framework utilizes the sample efficiency of proximal policy optimization to explore the exponentially large space of possible assignment of the bitwidths to the layers. We show how ReLeQ can balance speed and quality, and provide a heterogeneous bitwidth assignment for quantization of a large variety of deep networks with minimal accuracy loss (≤ 0.3% loss) while minimizing the computation and storage costs. With these DNNs, ReLeQ enables conventional hardware and custom DNN accelerator to achieve 2.2× speedup over 8-bit execution.