Leveraging MLC STT-RAM for energy-efficient CNN training

Leveraging MLC STT-RAM for energy-efficient CNN training
复制标题

DOI:
10.1145/3240302.3240422
复制
发表时间:
2018-10
期刊:
Proceedings of the International Symposium on Memory Systems
影响因子:
--
通讯作者:
Hengyu Zhao;Jishen Zhao
Hengyu Zhao;Jishen Zhao
中科院分区:
其他
文献类型:
--
作者:
Hengyu Zhao;Jishen Zhao

文献摘要

相似文献

图形处理单元(GPU)因其良好的计算能力而被广泛用于卷积神经网络(CNN)的训练。然而,随着训练模型的日益庞大和深入,GPU的内存容量、带宽和能量正成为关键的系统瓶颈。本文提出了一种节能的GPU内存管理方案,采用MLC STT-RAM作为GPU内存,以适应图像分类训练的工作量。我们提出了一种数据重映射方案,该方案利用了MLC STT-RAM单元中软、硬位之间的访问延迟和能量的不对称性以及图像分类训练负载中的内存访问特性。此外,我们的设计能够(I)通过利用训练数据中的位级相似性来实现节能的存储器访问,以及(Ii)最佳的特征图编码来压缩特征图中连续的0。我们的设计将VGG-19和AlexNet的训练时间、GPU内存访问能量和容量利用率分别降低了76%和70%、45%和40%、26.9%和26%。
Graphics Processing Units (GPUs) are extensively used in training of convolutional neural networks (CNNs) due to their promising compute capability. However, GPU memory capacity, bandwidth, and energy are becoming critical system bottlenecks with increasingly larger and deeper training models. This paper proposes an energy-efficient GPU memory management scheme by employing MLC STT-RAM as GPU memory to accommodate the image classification training workloads. We propose a data remapping scheme that exploits the asymmetry access latency and energy across soft and hard bits in MLC STT-RAM cells and the memory access characteristics in image classification training workloads. Furthermore, our design enables (i) energy-efficient memory access by leveraging bit-level similarity in training data and (ii) optimal feature map encoding to compress the contiguous 0s in feature maps. Our design reduces VGG-19 and AlexNet training time, GPU memory access energy and capacity utilization by 76% and 70%, 45% and 40%, 26.9% and 26%, respectively.