An SRAM-Based Multibit In-Memory Matrix-Vector Multiplier With a Precision That Scales Linearly in Area, Time, and Power

An SRAM-Based Multibit In-Memory Matrix-Vector Multiplier With a Precision That Scales Linearly in Area, Time, and Power
复制标题

基于 SRAM 的多位内存矩阵矢量乘法器,其精度可在面积、时间和功耗上线性缩放

DOI:
--
复制
发表时间:
2021
影响因子:
2.8
通讯作者:
E. Eleftheriou
E. Eleftheriou
中科院分区:
工程技术2区
文献类型:
--
作者:
R. Khaddam;P. Francese;L. Benini;E. Eleftheriou

文献摘要

参考文献

被引文献

相似文献

本文提出了一种用于内存计算的新型交错式开关电容和基于静态随机存取存储器(SRAM)的多位矩阵 - 向量乘累加引擎。其工作原理是首先使用由\(n + 1\)个等尺寸级构成的流水线数模转换器将存储在SRAM中的\(n\)位权重转换为成比例的电压。然后,一个开关电容级将这些电压与\(m\)位数字输入激活值相乘。最后,通过电荷共享将对应不同乘法结果的输出电压沿一列累加。利用我们提出的架构,所需的电路面积、计算时间和功耗与输入和权重的位分辨率呈线性关系。给出了电容和开关能耗的解析公式。此外,研究了制造失配对模拟计算精度的影响。描述了整个系统架构,并通过在14纳米工艺下的完整宏实现研究证明了其可行性,详细说明了面积和能耗以及总体延迟。最后,介绍了一个在14纳米工艺下的\(128×2048\)、\(6\)位权重和\(6\)位输入有符号矩阵 - 向量乘法加速器系统的具体设计,该系统在\(0.8V\)的标称电源电压下,运行速度为\(2.43\)万亿次操作每秒(TOP/s),效率为\(16.94\)万亿次操作每秒每瓦(TOP/s/W)。如果在指标中考虑操作数的精度,那么效率变为\(609.7\)万亿次操作每秒每瓦(TOP/s/W)。
A novel interleaved switched-capacitor and SRAM-based multibit matrix-vector multiply-accumulate engine for in-memory computing is presented. Its operation principle is based on first converting an SRAM-stored n-bit weight into a proportional voltage using a pipeline D/A converter built from $n+1$ equally sized stages. A switched-capacitor stage then multiplies these voltages with an m-bit digital input activation. Finally, the output voltages that correspond to the different multiplication results are accumulated along one column by means of charge-sharing. With our proposed architecture, the required circuit area, computation time, and power consumption scale linearly versus the bit resolution of both the inputs and the weights. Analytical formulas are presented for the energy consumption in both capacitors and switches. Moreover, the impact of fabrication mismatch on analog computation accuracy is examined. The full system architecture is described, and the feasibility is demonstrated, via a full macroimplementation study in 14 nm, detailing area and energy consumption, as well as the overall latency. Finally, a specific design of a $128 imes 2048,,6$ -bit weight and 6-bit input signed matrix-vector multiplication accelerator system in 14 nm is presented, which runs at 2.43 TOP/s at an efficiency of 16.94 TOP/s/W, while using the nominal supply voltage of 0.8 V. If the operands’ precision is considered in the metric, then the efficiency becomes 609.7 TOP/s/W.
DOI: 10.1109/jssc.2020.2992886
发表时间: 2020-07-01
影响因子: 5.4
作者:
Jiang, Zhewei;Yin, Shihui;Seok, Mingoo
通讯作者: Seok, Mingoo