Fundamental Limits on the Precision of In-memory Architectures

Fundamental Limits on the Precision of In-memory Architectures
复制标题

内存架构精度的基本限制

DOI:
--
复制
发表时间:
2020
期刊:
2020 IEEE/ACM International Conference On Computer Aided Design (ICCAD)
影响因子:
--
通讯作者:
Naresh R Shanbhag
Naresh R Shanbhag
中科院分区:
--
文献类型:
--
作者:
Sujan Kumar Gonugondla;Charbel Sakr;Hassan Dbouk;Naresh R Shanbhag

文献摘要

被引文献

相似文献

本文得到了内存计算架构(IMC)的计算精度的基本限制。定义了IMC的各种计算SNR指标,并分析了它们之间的相互关系,以表明IMC的准确性从根本上受到其模拟核心的计算SNR(SNRa)的限制,并且需要为最终输出SNR SNRT → SNRa适当地分配激活、权重和输出精度。最小精度准则(MPC)的建议,以尽量减少输出,从而列模数转换器(ADC)的精度。研究了电荷求和(QS)计算模型及其相关的IMC QS-Arch,获得了计算SNR、最小ADC精度、能量和延迟的解析模型。QS-Arch的计算SNR模型在65 nm CMOS工艺中通过Monte Carlo模拟进行了验证。采用这些模型,上界的QS-Arch-为基础的IMC采用512行SRAM阵列的SNRa,它表明,QS-Arch的能量成本降低了3.3倍,每6 dB的SNRa下降,最大可实现的SNRa降低技术缩放,而在相同的SNRa的能量成本增加。这些模型还表明,由于电压净空限幅,点积维度N存在上限,并且SNRa每下降3 dB,该上限可以加倍。
This paper obtains the fundamental limits on the computational precision of in-memory computing architectures (IMCs). Various compute SNR metrics for IMCs are defined and their interrelationships analyzed to show that the accuracy of IMCs is fundamentally limited by the compute SNR (SNRa) of its analog core, and that activation, weight and output precision needs to be assigned appropriately for the final output SNR SNRT → SNRa. The minimum precision criterion (MPC) is proposed to minimize the output and hence the column analog-to-digital converter (ADC) precision. The charge summing (QS) compute model and its associated IMC QS-Arch are studied to obtain analytical models for its compute SNR, minimum ADC precision, energy and latency. Compute SNR models of QS-Arch are validated via Monte Carlo simulations in a 65 nm CMOS process. Employing these models, upper bounds on SNRa of a QS-Arch-based IMC employing a 512 row SRAM array are obtained and it is shown that QS-Arch's energy cost reduces by 3.3× for every 6 dB drop in SNRa, and that the maximum achievable SNRa reduces with technology scaling while the energy cost at the same SNRa increases. These models also indicate the existence of an upper bound on the dot product dimension N due to voltage headroom clipping, and this bound can be doubled for every 3 dB drop in SNRa.