Exploring Compute-in-Memory Architecture Granularity for Structured Pruning of Neural Networks

Exploring Compute-in-Memory Architecture Granularity for Structured Pruning of Neural Networks
复制标题

DOI:
10.1109/jetcas.2022.3227471
复制
发表时间:
2022-12
影响因子:
4.6
通讯作者:
F. Meng;Xinxin Wang;Ziyu Wang;Eric Lee;Wei D. Lu
F. Meng;Xinxin Wang;Ziyu Wang;Eric Lee;Wei D. Lu
中科院分区:
工程技术2区
文献类型:
--
作者:
F. Meng;Xinxin Wang;Ziyu Wang;Eric Lee;Wei D. Lu

文献摘要

相似文献

使用电阻式随机存储器(RRAM)交叉棒实现的内存计算(CIM)是深度神经网络(DNN)加速的一种很有前途的方法。随着DNN规模的持续增长,有限的片上权重存储已成为CIM实现的挑战。剪枝可以减小网络规模,但非结构化剪枝与CIM不兼容,而结构化剪枝会导致更高的神经网络精度下降。在这项工作中,我们系统地评估如何结构化修剪可以有效地实现CIM系统。我们表明,通过利用CIM操作中固有的计算粒度,细粒度的结构化修剪可以支持提高精度和最小的硬件成本。我们讨论了在一个实际的系统中的硬件实现和预期的性能的准确性,能源和有效的吞吐量。与所提出的方法,压缩比高达11.1(即9%的权重剩余),可以实现只有0.6%的精度下降,在硬件设计中的硬件开销最小。
Compute-in-Memory (CIM) implemented with Resistive-Random-Access-Memory (RRAM) crossbars is a promising approach for Deep Neural Network (DNN) acceleration. As the DNN size continues to grow, the finite on-chip weight storage has become a challenge for CIM implementations. Pruning can reduce network size, but unstructured pruning is not compatible with CIM, while structured pruning leads to higher neural network accuracy drop. In this work we systematically evaluate how structured pruning can be efficiently implemented in CIM systems. We show that by utilizing the inherent computational granularity in CIM operations, fine-grained structured pruning can be supported with improved accuracy and minimal hardware cost. We discuss the hardware implementation in a practical system and the expected performance in terms of accuracy, energy and effective throughput. With the proposed approach, compression ratio up to 11.1 (i.e. 9% weights remaining) can be achieved with only 0.6% accuracy drop with minimal hardware overhead in the hardware design.