Optimization of DRAM based PIM Architecture for Energy-Efficient Deep Neural Network Training

Optimization of DRAM based PIM Architecture for Energy-Efficient Deep Neural Network Training
复制标题

基于 DRAM 的 PIM 架构优化,用于节能深度神经网络训练

DOI:
10.1109/iscas48785.2022.9937832
复制
发表时间:
2022
期刊:
2022 IEEE International Symposium on Circuits and Systems (ISCAS)
影响因子:
--
通讯作者:
N. Wehn
N. Wehn
中科院分区:
--
文献类型:
--
作者:
C. Sudarshan;Mohammad Hassani Sadi;C. Weis;N. Wehn

文献摘要

参考文献

被引文献

相似文献

深度神经网络(DNN)训练消耗高能量。另一方面,部署在边缘设备上的DNN需要非常高的能效。在这种背景下,内存处理(PIM)是一种新兴的计算范式,它弥合了内存计算的差距,以提高能源效率。DRAM是一种这样的存储器类型,用于设计用于DNN训练的节能PIM架构。为DNN训练设计的DRAM-PIM架构的主要问题之一是存储器阵列和PIM计算单元之间的库内的大量内部数据访问(例如,比推断多51%)。与计算单元相比,最先进的DRAM PIM架构中的这些内部数据访问消耗非常高的能量。因此,重要的是减少DRAM存储体内的内部数据访问能量,以进一步提高DRAMPIM架构的能量效率。我们提出了三种新的优化,共同降低内部数据访问能量高达81.54%。我们的第一个优化修改的银行数据访问电路,使部分访问的数据,而不是传统的固定粒度的访问,从而利用在训练过程中可用的稀疏性。第二个优化是在DRAM组内具有专用的低能量区域,其具有全局线的低电容负载和较短的数据移动。最后,我们提出了一种12位高动态范围浮点格式,称为TinyFloat,与IEEE 754半精度和单精度相比,减少了20%的数据访问能量总数。
Deep Neural Network (DNN) training consumes high-energy. On the other hand, DNNs deployed on edge devices demand very high-energy efficiency. In this context, Processing-in-Memory (PIM) is an emerging compute paradigm that bridges the memory-computation gap to improve the energy-efficiency. DRAMs are one such memory type employed for designing energy-efficient PIM architectures for DNN training. One of the major issues of DRAM-PIM architectures designed for DNN training is the high number of internal data accesses within a bank between the memory arrays and the PIM computation units (e.g. 51% more than inference). These internal data accesses in the state-of-the-art DRAM PIM architectures consume very high energy compared to computation units. Hence, it is important to reduce the internal data access energy within the DRAM bank for further improving the energy efficiency of DRAMPIM architectures. We present three novel optimizations that together reduce the internal data access energy up to 81.54%. Our first optimization modifies the bank data access circuit to enable partial accesses of data instead of the conventional fixed granularity accesses, thereby exploiting the available sparsity during training. The second optimization is to have a dedicated low-energy region within the DRAM bank that has low capacitive load of global wires and shorter data movement. Finally, we propose a 12-bit high dynamic range floating-point format called TinyFloat that reduces the total number of data access energy by 20% compared to IEEE 754 half and single precision.
DRAMSpec:高级 DRAM 时序、功耗和面积探索工具
DOI: 10.1007/s10766-016-0473-y
发表时间: 2016
影响因子: 1.5
作者:
C. Weis;A. Mutaal;O. Naji;M. Jung;A. Hansson;N. Wehn
通讯作者: N. Wehn
DrAcc:基于 DRAM 的准确 CNN 推理加速器
DOI: 10.1109/dac.2018.8465866
发表时间: 2018
期刊: 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC
影响因子: --
作者:
Deng, Quan;Jiang, Lei;Zhang, Youtao;Zhang, Minxuan;Yang, Jun
通讯作者: Yang, Jun