AI hardware acceleration with analog memory: Microarchitectures for low energy at high speed

AI hardware acceleration with analog memory: Microarchitectures for low energy at high speed
复制标题

DOI:
10.1147/jrd.2019.2934050
复制
发表时间:
2019-11-01
影响因子:
1.3
通讯作者:
Burr, G. W.
Burr, G. W.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Chang, H-Y;Narayanan, P.;Burr, G. W.

文献摘要

被引文献

相似文献

在本文中,我们提出了在模拟存储器交叉阵列中实现的多层深度神经网络(DNN)的创新微架构设计。数据在阵列之间以完全并行的方式传输,无需显式模数转换器。采用基于源跟随器的读出、阵列分段和按持续时间发送等设计思想来提高电路效率。使用 90 nm 技术节点中完整 CMOS 设计的电路仿真来定量分析 DNN 训练和推理的执行能量和吞吐量。我们发现,我们当前的设计可以实现高达 12-14 TOPs/s/W 的训练能效,而预计的扩展设计可以实现高达 250 TOPs/s/W。讨论了实现模拟人工智能系统的关键挑战。
In this article, we present innovative microarchitectural designs for multilayer deep neural networks (DNNs) implemented in crossbar arrays of analog memories. Data is transferred in a fully parallel manner between arrays without explicit analog-to-digital converters. Design ideas including source follower-based readout, array segmentation, and transmit-by-duration are adopted to improve the circuit efficiency. The execution energy and throughput, for both DNN training and inference, are analyzed quantitatively using circuit simulations of a full CMOS design in the 90-nm technology node. We find that our current design could achieve up to 12-14 TOPs/s/W energy efficiency for training, while a projected scaled design could achieve up to 250 TOPs/s/W. Key challenges in realizing analog AI systems are discussed.