Processing-in-Memory for Energy-Efficient Neural Network Training: A Heterogeneous Approach

Processing-in-Memory for Energy-Efficient Neural Network Training: A Heterogeneous Approach
复制标题

DOI:
10.1109/micro.2018.00059
复制
发表时间:
2018-10
期刊:
2018 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Jiawen Liu;Hengyu Zhao;Matheus A. Ogleari;Dong Li;Jishen Zhao
Jiawen Liu;Hengyu Zhao;Matheus A. Ogleari;Dong Li;Jishen Zhao
中科院分区:
其他
文献类型:
--
作者:
Jiawen Liu;Hengyu Zhao;Matheus A. Ogleari;Dong Li;Jishen Zhao

文献摘要

被引文献

相似文献

神经网络(NN)已被广泛应用于图像分类、语音识别、目标检测和计算机视觉等领域。然而,训练神经网络-特别是深度神经网络(DNN)-可能会消耗能量和时间,因为处理器和存储器之间的频繁数据移动。此外,训练涉及具有各种计算和存储器访问特征的大量细粒度操作。利用这种不同操作的高度并行性是具有挑战性的。为了解决这些问题,我们提出了一种异构内存处理(PIM)系统的软件/硬件协同设计。我们的硬件设计采用了数百个固定功能的算术单元和基于ARM的可编程内核的逻辑层上的3D芯片堆叠存储器,形成一个异构的PIM架构附加到CPU。我们的软件设计提供了一个编程模型和一个运行时系统,可以跨CPU和异构PIM提供的计算资源对各种NN训练操作进行编程、卸载和调度。通过扩展OpenCL编程模型和采用硬件异构感知运行时系统,我们实现了跨各种异构硬件的高程序可移植性和轻松的程序维护,优化系统能源效率,并提高硬件利用率。
Neural networks (NNs) have been adopted in a wide range of application domains, such as image classification, speech recognition, object detection, and computer vision. However, training NNs – especially deep neural networks (DNNs) – can be energy and time consuming, because of frequent data movement between processor and memory. Furthermore, training involves massive fine-grained operations with various computation and memory access characteristics. Exploiting high parallelism with such diverse operations is challenging. To address these challenges, we propose a software/hardware co-design of heterogeneous processing-in-memory (PIM) system. Our hardware design incorporates hundreds of fix-function arithmetic units and ARM-based programmable cores on the logic layer of a 3D die-stacked memory to form a heterogeneous PIM architecture attached to CPU. Our software design offers a programming model and a runtime system that program, offload, and schedule various NN training operations across compute resources provided by CPU and heterogeneous PIM. By extending the OpenCL programming model and employing a hardware heterogeneity-aware runtime system, we enable high program portability and easy program maintenance across various heterogeneous hardware, optimize system energy efficiency, and improve hardware utilization.