Enabling NVM-based deep learning acceleration using nonuniform data quantization: work-in-progress

Enabling NVM-based deep learning acceleration using nonuniform data quantization: work-in-progress
复制标题

DOI:
10.1145/3125501.3125516
复制
发表时间:
2017-10
期刊:
Proceedings of the 2017 International Conference on Compilers, Architectures and Synthesis for Embedded Systems Companion
影响因子:
--
通讯作者:
Hao Yan;Ethan C. Ahn;Lide Duan
Hao Yan;Ethan C. Ahn;Lide Duan
中科院分区:
其他
文献类型:
--
作者:
Hao Yan;Ethan C. Ahn;Lide Duan

文献摘要

相似文献

除了使用神经网络(NN)计算的协同处理器(例如GPU),还利用非挥发性记忆的独特特征(NVM),包括RRAM,相变内存(PCM)和STT-MRAM,以加速NN算法NN算法已经进行了广泛的研究。在这种方法中,输入数据和突触权重用单词线电压和电池电阻表示,结果位线电流表示计算结果。但是,NVM细胞中的电阻水平有限大大降低了算法数据精度,因此显着降低了模型推理精度。通过观察到,传统的,均匀生成的数据量化点对模型并不重要的动机,我们提出了一种非均匀的数据量化方案,以更好地表示NVM细胞中的模型并最大程度地减少推理准确性损失。我们的实验结果表明,所提出的方案可以实现高度准确的深度学习模型推断,其低至4位用于突触体重表示。这有效地使NVM较少具有细胞电阻水平(例如STT-MRAM)执行NN计算,并且还会在性能,能量和存储器存储方面带来其他好处。
Apart from employing a co-processor (e.g., GPU) for neural network (NN) computation, utilizing the unique characteristics of nonvolatile memories (NVM), including RRAM, phase change memory (PCM), and STT-MRAM, to accelerate NN algorithms has been extensively studied. In such approaches, input data and synaptic weights are represented using word line voltages and cell resistance, with the resulting bit line current indicating the calculation result. However, the limited number of resistance levels in a NVM cell largely reduces the algorithm data precision, thus significantly lowering the model inference accuracy. Motivated by the observation that the conventional, uniformly generated data quantization points are not equally important to the model, we propose a nonuniform data quantization scheme to better represent the model in NVM cells and minimize the inference accuracy loss. Our experimental results show that the proposed scheme can achieve highly accurate deep learning model inference using as low as only 4 bits for synaptic weight representation. This effectively enables a NVM with few cell resistance levels (e.g., STT-MRAM) to perform NN calculation, and also results in additional benefits in performance, energy, and memory storage.