A 28-nm Compute SRAM With Bit-Serial Logic/Arithmetic Operations for Programmable In-Memory Vector Computing

A 28-nm Compute SRAM With Bit-Serial Logic/Arithmetic Operations for Programmable In-Memory Vector Computing
复制标题

DOI:
10.1109/jssc.2019.2939682
复制
发表时间:
2020-01
影响因子:
5.4
通讯作者:
Jingcheng Wang;Xiaowei Wang;Charles Eckert;Arun K. Subramaniyan;R. Das;D. Blaauw;D. Sylvester
Jingcheng Wang;Xiaowei Wang;Charles Eckert;Arun K. Subramaniyan;R. Das;D. Blaauw;D. Sylvester
中科院分区:
工程技术1区
文献类型:
--
作者:
Jingcheng Wang;Xiaowei Wang;Charles Eckert;Arun K. Subramaniyan;R. Das;D. Blaauw;D. Sylvester

文献摘要

被引文献

相似文献

本文提出了一种通用混合体内/接近内存的计算SRAM(CRAM),该混合物将8T可转座的位细胞与基于矢量的,基于矢量的,比特的内存算术算术相结合,可容纳各种位宽度,来自单个位宽度。到32或64位,以及一组完整的操作类型,包括整数和浮点添加,乘法和划分。这种方法提供了从神经网络到图形和信号处理的不断发展的软件算法所需的灵活性和可编程性。所提出的设计是在由28 nm CMO的小型物联网(IoT)处理器中实施的,该处理器由Cortex-M0 CPU和8个16 KB的CRAM BANK组成(总计128 kb)。该系统在1.1 V时可实现475 MHz的操作,并且在所有CRAMS活跃的情况下,在32位操作数上产生30 GOPS或1.4 Gflops。它可实现8位乘法的0.56 TOP/W的能源效率,在0.6 V和114 MHz时添加8位的效率为5.27 TOPS/W。
This article proposes a general-purpose hybrid in-/near-memory compute SRAM (CRAM) that combines an 8T transposable bit cell with vector-based, bit-serial in-memory arithmetic to accommodate a wide range of bit-widths, from single to 32 or 64 bits, as well as a complete set of operation types, including integer and floating-point addition, multiplication, and division. This approach provides the flexibility and programmability necessary for evolving software algorithms ranging from neural networks to graph and signal processing. The proposed design was implemented in a small Internet of Things (IoT) processor in the 28-nm CMOS consisting of a Cortex-M0 CPU and 8 CRAM banks of 16 kB each (128 kB total). The system achieves 475-MHz operation at 1.1 V and, with all CRAMs active, produces 30 GOPS or 1.4 GFLOPS on 32-bit operands. It achieves an energy efficiency of 0.56 TOPS/W for 8-bit multiplication and 5.27 TOPS/W for 8-bit addition at 0.6 V and 114 MHz.