aCortex: An Energy-Efficient Multipurpose Mixed-Signal Inference Accelerator

aCortex: An Energy-Efficient Multipurpose Mixed-Signal Inference Accelerator
复制标题

DOI:
10.1109/jxcdc.2020.2999581
复制
发表时间:
2020-06-01
影响因子:
2.4
通讯作者:
Strukov, Dmitri B.
Strukov, Dmitri B.
中科院分区:
其他
文献类型:
--
作者:
Bavandpour, Mohammad;Mahmoodi, Mohammad R.;Strukov, Dmitri B.

文献摘要

被引文献

相似文献

我们介绍了“aCortex”,这是一种非常节能、快速、紧凑和多功能的神经形态处理器体系结构,适用于加速各种神经网络推理模型。我们的处理器最重要的功能是一个可配置的混合信号计算阵列,它由向量乘矩阵乘法器(VMM)块组成,使用嵌入式非易失性存储器阵列来存储权重矩阵。用于数据转换和高压编程的模拟外围电路在大量VMM块之间共享,以促进不同类型神经网络层的紧凑和节能的模拟域VMM操作。ACortex的其他独特功能包括可配置的缓冲器和数据总线链、简单高效的指令集体系结构及其相应的多代理控制器、可编程量化范围以及定制的无刷新嵌入式动态随机存取存储器。采用内嵌NOR闪存的55 nm工艺设计了具有4位模拟计算精度的能量最优aCortex。它的物理性能是使用测试几个常见基准的单个电路元件和关键部件物理布局的实验数据来评估的,这些基准是用于图像分类的两个最先进的深度前馈网络Inception-VL和ResNet-152,以及用于语言翻译的谷歌深度递归网络GNTM。这些基准测试的系统级仿真结果分别显示了97、106和336top/J的能效,以及高达15top/S的计算吞吐量和0.27MB/mm(2)的存储效率。这种估计的性能结果与之前报道的基于不太成熟的积极扩展的阻性开关存储器的混合信号加速器相比是有利的。
We introduce "aCortex," an extremely energy-efficient, fast, compact, and versatile neuromorphic processor architecture suitable for the acceleration of a wide range of neural network inference models. The most important feature of our processor is a configurable mixed-signal computing array of vector-by-matrix multiplier (VMM) blocks utilizing embedded nonvolatile memory arrays for storing weight matrices. Analog peripheral circuitry for data conversion and high-voltage programming are shared among a large array of VMM blocks to facilitate compact and energy-efficient analog-domain VMM operation of different types of neural network layers. Other unique features of aCortex include configurable chain of buffers and data buses, simple and efficient instruction set architecture and its corresponding multiagent controller, programmable quantization range, and a customized refresh-free embedded dynamic random access memory. The energy-optimal aCortex with 4-bit analog computing precision was designed in a 55-nm process with embedded NOR flash memory. Its physical performance was evaluated using experimental data from testing individual circuit elements and physical layout of key components for several common benchmarks, namely, Inception-vl and ResNet-152, two state-of-the-art deep feedforward networks for image classification, and GNTM, Google's deep recurrent network for language translation. The system-level simulation results for these benchmarks show the energy efficiency of 97, 106, and 336 TOp/J, respectively, combined with up to 15 TOp/s computing throughput and 0.27-MB/mm(2) storage efficiency. Such estimated performance results compare favorably with those of previously reported mixed-signal accelerators based on much less mature aggressively scaled resistive switching memories.