CASH-RAM: Enabling In-Memory Computations for Edge Inference Using Charge Accumulation and Sharing in Standard 8T-SRAM Arrays

CASH-RAM: Enabling In-Memory Computations for Edge Inference Using Charge Accumulation and Sharing in Standard 8T-SRAM Arrays
复制标题

DOI:
10.1109/jetcas.2020.3014250
复制
发表时间:
2020-09-01
影响因子:
4.6
通讯作者:
Roy, Kaushik
Roy, Kaushik
中科院分区:
工程技术2区
文献类型:
--
作者:
Agrawal, Amogh;Kosta, Adarsh;Roy, Kaushik

文献摘要

被引文献

相似文献

机器学习(ML)工作负载是内存和计算密集型的,在传统计算系统上运行时会消耗大量电力,从而将其实现限制在大型数据中心。将大量数据从边缘设备传输到数据中心不仅能源昂贵,而且有时在安全关键型应用中也不可取。因此,需要构建特定于域的硬件原语,用于边缘处的节能ML处理。一种这样的方法-内存计算,通过直接计算存储数据的地方,消除了内存和计算单元之间频繁和不必要的数据传输。然而,计算的模拟性质引入了非理想性,这降低了神经网络的整体精度。在本文中,我们提出了一个在内存中的计算原语加速点产品在标准的8 T-SRAM高速缓存,使用电荷共享。位线和源极线的固有寄生电容用于累积模拟电压,其可以被感测为近似点积。电荷共享方法涉及自补偿技术,其减少非理想性的影响,从而减少误差。我们的研究结果表明,使用所提出的补偿方法,精度下降是在1%和5%的基线精度,分别为MNIST和CIFAR-10数据集,与能量延迟产品的改进38\times $比标准冯诺依曼计算系统。我们认为,这项工作可以与现有的缓解技术,如再培训的方法,以进一步提高系统的性能。
Machine Learning (ML) workloads being memory- and compute-intensive, consume large amounts of power running on conventional computing systems, restricting their implementations to large-scale data centers. Transferring large amounts of data from the edge devices to the data centers is not only energy expensive, but sometimes undesirable in security-critical applications. Thus, there is a need for building domain-specific hardware primitives for energy-efficient ML processing at the edge. One such approach - in-memory computing, eliminates frequent and unnecessary data-transfers between the memory and the compute units, by directly computing the data where it is stored. However, the analog nature of computations introduces non-idealities, which degrades the overall accuracy of neural networks. In this paper, we propose an in-memory computing primitive for accelerating dot-products within standard 8T-SRAM caches, using charge-sharing. The inherent parasitic capacitance of the bitlines and sourcelines is used for accumulating analog voltages, which can be sensed for an approximate dot product. The charge sharing approach involves a self-compensation technique which reduces the effects of non-idealities, thereby reducing the errors. Our results for ternary weight neural networks show that using the proposed compensation approaches, the accuracy degradation is within 1% and 5% of the baseline accuracy, for the MNIST and CIFAR-10 dataset, respectively, with an energy-delay product improvement of $38\times $ over the standard von-Neumann computing system. We believe that this work can be used in conjunction with existing mitigation techniques, such as re-training approaches, to further enhance system performance.