Tiny but Accurate: A Pruned, Quantized and Optimized Memristor Crossbar Framework for Ultra Efficient DNN Implementation

Tiny but Accurate: A Pruned, Quantized and Optimized Memristor Crossbar Framework for Ultra Efficient DNN Implementation
复制标题

DOI:
10.1109/asp-dac47756.2020.9045658
复制
发表时间:
2019-08
期刊:
2020 25th Asia and South Pacific Design Automation Conference (ASP-DAC)
影响因子:
--
通讯作者:
Xiaolong Ma;Geng Yuan;Sheng Lin;Caiwen Ding;Fuxun Yu;Tao Liu;Wujie Wen;Xiang Chen;Yanzhi Wang
Xiaolong Ma;Geng Yuan;Sheng Lin;Caiwen Ding;Fuxun Yu;Tao Liu;Wujie Wen;Xiang Chen;Yanzhi Wang
中科院分区:
其他
文献类型:
--
作者:
Xiaolong Ma;Geng Yuan;Sheng Lin;Caiwen Ding;Fuxun Yu;Tao Liu;Wujie Wen;Xiang Chen;Yanzhi Wang

文献摘要

相似文献

忆阻器交叉杆阵列已经成为DNN应用的内在合适的矩阵计算和低功耗加速框架。基于忆阻器的权值修剪和量化等技术已经被研究。然而,上述技术的高精度解决方案仍有待解开。在本文中,我们提出了一种基于忆阻器的DNN框架,它结合了结构化的权重修剪和量化,通过结合ADMM算法,以获得更好的修剪和量化性能。我们还发现了ADMM解决方案的非最优权修剪和未使用的数据路径结构化修剪模型。我们设计了一个软硬件协同优化框架,其中包含第一个提出的网络净化和未使用的路径删除算法,针对后处理的结构化修剪模型后ADMM步骤。通过将忆阻器硬件约束纳入我们的整个框架,我们以最小的精度损失实现了极高的压缩率。对于量化结构化修剪模型,我们的框架在将权重量化为8位忆阻器权重表示后几乎没有精度损失。我们在匿名链接https://bit.ly/2VnMUy0上分享我们的模型。
The memristor crossbar array has emerged as an intrinsically suitable matrix computation and low-power acceleration framework for DNN applications. Many techniques such as memristor-based weight pruning and memristor-based quantization have been studied. However, the high accuracy solution for the above techniques is still waiting for unraveling. In this paper, we propose a memristor-based DNN framework which combines both structured weight pruning and quantization by incorporating ADMM algorithm for better pruning and quantization performance. We also discover the non-optimality of the ADMM solution in weight pruning and the unused data path in a structured pruned model. We design a software-hardware co-optimization framework which contains the first proposed Network Purification and Unused Path Removal algorithms targeting on post-processing a structured pruned model after ADMM steps. By taking memristor hardware constraints into our whole framework, we achieve extreme high compression rate with minimum accuracy loss. For quantizing structured pruned model, our framework achieves nearly no accuracy loss after quantizing weights to 8-bit memristor weight representation. We share our models at anonymous link https://bit.ly/2VnMUy0.