C3SRAM: An In-Memory-Computing SRAM Macro Based on Robust Capacitive Coupling Computing Mechanism

C3SRAM: An In-Memory-Computing SRAM Macro Based on Robust Capacitive Coupling Computing Mechanism
复制标题

DOI:
10.1109/jssc.2020.2992886
复制
发表时间:
2020-07-01
影响因子:
5.4
通讯作者:
Seok, Mingoo
Seok, Mingoo
中科院分区:
工程技术1区
文献类型:
--
作者:
Jiang, Zhewei;Yin, Shihui;Seok, Mingoo

文献摘要

被引文献

相似文献

本文介绍了C3SRAM,一种内存计算型SRAM宏。该宏是一个SRAM模块,其电路嵌入在位单元和外设中,用于对具有二值化权重和激活值的神经网络进行硬件加速。该宏利用模拟混合信号(AMS)电容耦合计算来评估二值神经网络的主要计算,即二值乘累加操作。无需逐行访问存储的权重,该宏同时激活所有行,并通过电容分压在读位线节点形成模拟电压。每列有一个模数转换器(ADC),该宏在单个周期内实现完全并行的向量 - 矩阵乘法。该宏所支持的网络类型及其采用的计算机制由AMS计算所需的鲁棒性和容错性决定。C3SRAM宏在65纳米CMOS工艺中进行了原型制作。它展示出672万亿次运算每秒每瓦的能效以及16380亿次运算每秒(20.2万亿次运算每秒每平方毫米)的速度,与执行相同操作的传统数字基准相比,其能量延迟积提高了3975倍。对于MNIST数据集,该宏达到98.3%的准确率,对于CIFAR - 10数据集达到85.5%的准确率,在能效和推理准确率的权衡方面,处于内存计算领域的领先水平。
This article presents C3SRAM, an in-memory-computing SRAM macro. The macro is an SRAM module with the circuits embedded in bitcells and peripherals to perform hardware acceleration for neural networks with binarized weights and activations. The macro utilizes analog-mixed-signal (AMS) capacitive-coupling computing to evaluate the main computations of binary neural networks, binary-multiply-and-accumulate operations. Without the need to access the stored weights by individual row, the macro asserts all its rows simultaneously and forms an analog voltage at the read bitline node through capacitive voltage division. With one analog-to-digital converter (ADC) per column, the macro realizes fully parallel vector-matrix multiplication in a single cycle. The network type that the macro supports and the computing mechanism it utilizes are determined by the robustness and error tolerance necessary in AMS computing. The C3SRAM macro is prototyped in a 65-nm CMOS. It demonstrates an energy efficiency of 672 TOPS/W and a speed of 1638 GOPS (20.2 TOPS/mm(2)), achieving 3975x better energy-delay product than the conventional digital baseline performing the same operation. The macro achieves 98.3% accuracy for MNIST and 85.5% for CIFAR-10, which is among the best in-memory computing works in terms of energy efficiency and inference accuracy tradeoff.