AGNI: In-Situ, Iso-Latency Stochastic-to-Binary Number Conversion for In-DRAM Deep Learning

AGNI: In-Situ, Iso-Latency Stochastic-to-Binary Number Conversion for In-DRAM Deep Learning
复制标题

DOI:
10.1109/isqed57927.2023.10129301
复制
发表时间:
2023-02
期刊:
2023 24th International Symposium on Quality Electronic Design (ISQED)
影响因子:
--
通讯作者:
Supreeth Mysore Shivanandamurthy;Sairam Sri Vatsavai;Ishan G. Thakkar;S. A. Salehi
Supreeth Mysore Shivanandamurthy;Sairam Sri Vatsavai;Ishan G. Thakkar;S. A. Salehi
中科院分区:
其他
文献类型:
--
作者:
Supreeth Mysore Shivanandamurthy;Sairam Sri Vatsavai;Ishan G. Thakkar;S. A. Salehi

文献摘要

相似文献

近年来,基于DRAM的存储器中处理(PIM)加速器领域的研究活动迅速增加,其中通过最小程度地改变DRAM外围设备的固有结构来利用DRAM的模拟计算能力以加速各种以数据为中心的应用。几个基于DRAM的PIM加速器卷积神经网络(CNN)也有报道。其中,利用DRAM内随机算术的加速器在处理延迟和吞吐量方面显示出多方面的改进,这是由于随机算术将乘法转换为简单的逐位逻辑AND运算的能力。然而,使用DRAM中的随机算法进行CNN加速需要频繁的随机到二进制数的转换。为此,现有技术采用基于全加法器或基于串行计数器的DRAM内电路。这些电路消耗大面积并且引起长延迟。它们的DRAM内实现还需要对DRAM外围设备进行大量修改,这大大减少了在这些加速器中使用随机算法的好处。为了解决这些缺点,本文提出了一种新的基板在DRAM随机到二进制数转换称为AGNI。AGNI使用传输晶体管、电容器、编码器和电荷泵对DRAM外围设备进行了微小的修改,并将读出放大器重新用作电压比较器,以实现不同大小的输入统计操作数的原位二进制转换。基于详细的SPICE模拟(https://github.com/uky-UCAT/AGNI_SPICE.git),我们的评估表明,与之前的两个DRAM随机到二进制转换电路相比,AGNI可以实现至少8倍的面积节省,至少28倍的能量延迟乘积(EDP)和至少21倍的面积延迟。这些电路级优势被证明可以在系统级传播,在四个深度CNN模型中实现至少3.9倍的性能增益。
Recent years have seen a rapid increase in research activity in the field of DRAM-based Processing-In-Memory (PIM) accelerators, where the analog computing capability of DRAM is employed by minimally changing the inherent structure of DRAM peripherals to accelerate various data-centric applications. Several DRAM-based PIM accelerators for Convolutional Neural Networks (CNNs) have also been reported. Among these, the accelerators leveraging in-DRAM stochastic arithmetic have shown manifold improvements in processing latency and throughput, due to the ability of stochastic arithmetic to convert multiplications into simple bit-wise logical AND operations. However, the use of in-DRAM stochastic arithmetic for CNN acceleration requires frequent stochastic to binary number conversions. For that, prior works employ full adder-based or serial counter-based in-DRAM circuits. These circuits consume large area and incur long latency. Their in-DRAM implementations also require heavy modifications in DRAM peripherals, which significantly diminishes the benefits of using stochastic arithmetic in these accelerators. To address these shortcomings, this paper presents a new substrate for in-DRAM stochastic-to-binary number conversion called AGNI. AGNI makes minor modifications in DRAM peripherals using pass transistors, capacitors, encoders, and charge pumps, and re-purposes the sense amplifiers as voltage comparators, to enable in-situ binary conversion of input statistic operands of different sizes with iso latency. Our evaluations, based on detailed SPICE simulations (https://github.com/uky-UCAT/AGNI_SPICE.git), show that AGNI can achieve savings of at least 8× in area, at least 28× energy-delay product (EDP), and at least 21 in area × latency, compared to two in-DRAM stochastic-to-binary conversion circuits from prior works. These circuit-level benefits are demonstrated to propagate at the system-level to achieve at least 3.9× gain in performance across four deep CNN models.