MF-Net: Compute-In-Memory SRAM for Multibit Precision Inference Using Memory-Immersed Data Conversion and Multiplication-Free Operators

MF-Net: Compute-In-Memory SRAM for Multibit Precision Inference Using Memory-Immersed Data Conversion and Multiplication-Free Operators
复制标题

DOI:
10.1109/tcsi.2021.3064033
复制
发表时间:
2021-01
期刊:
IEEE Transactions on Circuits and Systems I: Regular Papers
影响因子:
--
通讯作者:
Shamma Nasrin;Diaa Badawi;A. Cetin;Wilfred Gomes;A. Trivedi
Shamma Nasrin;Diaa Badawi;A. Cetin;Wilfred Gomes;A. Trivedi
中科院分区:
其他
文献类型:
--
作者:
Shamma Nasrin;Diaa Badawi;A. Cetin;Wilfred Gomes;A. Trivedi

文献摘要

被引文献

相似文献

我们提出了一种用于计算深神经网络的内存推理的方法(DNN)。 RAM行/列,需要高精度类似于数字的转换器(ADC),对多位精度权重的支持有限,矢量尺度平行性无缝地扩展到多位精确的权重,它不需要DACS,它很容易扩展到较高的矢量范围。 SRAM数组的位线由于SA-ADC的主流区域的电容性DAC,因此DAC的电容性DAC是通过利用SRAM阵列的固有寄生虫来实现的,因此SRAM SA-ADC内的降低面积为62 $ SRAM Macro。 M CMOS SRAM宏需要4位ADC,需要较低的ADC精度的SRAM宏,但是我们还评估了MNIST NECTICE和CIFARES的组合使用从总计的85%的运算量为98.6%的MNIST,CIFAR10的乘法数量为98.6%,CIFAR100的运算率为90.2%,因为CIFAR100的大多数操作是基于提议的SRAM MACROS,我们的Compute-In-MaCros均受范围。
We propose a co-design approach for compute-in-memory inference for deep neural networks (DNN). We use multiplication-free function approximators based on $\ell _{1}$ norm along with a co-adapted processing array and compute flow. Using the approach, we overcame many deficiencies in the current art of in-SRAM DNN processing such as the need for digital-to-analog converters (DACs) at each operating SRAM row/column, the need for high precision analog-to-digital converters (ADCs), limited support for multi-bit precision weights, and limited vector-scale parallelism. Our co-adapted implementation seamlessly extends to multi-bit precision weights, it doesn’t require DACs, and it easily extends to higher vector-scale parallelism. We also propose an SRAM-immersed successive approximation ADC (SA-ADC), where we exploit the parasitic capacitance of bit lines of SRAM array as a capacitive DAC. Since the dominant area overhead in SA-ADC comes due to its capacitive DAC, by exploiting the intrinsic parasitic of SRAM array, our approach allows low area implementation of within-SRAM SA-ADC. Our $8\times 62$ SRAM macro, which requires a 5-bit ADC, achieves ~105 tera operations per second per Watt (TOPS/W) with 8-bit input/weight processing at 45 nm CMOS. Our $8\times 30$ SRAM macro, which requires a 4-bit ADC, achieves ~84 TOPS/W. SRAM macros that require lower ADC precision are more tolerant of process variability, however, have lower TOPS/W as well. We evaluated the accuracy and performance of our proposed network for MNIST, CIFAR10, and CIFAR100 datasets. We chose a network configuration which adaptively mixes multiplication-free and regular operators. The network configurations utilize the multiplication-free operator for more than 85% operations from the total. The selected configurations are 98.6% accurate for MNIST, 90.2% for CIFAR10, and 66.9% for CIFAR100. Since most of the operations in the considered configurations are based on proposed SRAM macros, our compute-in-memory’s efficiency benefits broadly translate to the system-level.