DeepNVM++: Cross-Layer Modeling and Optimization Framework of Nonvolatile Memories for Deep Learning

DeepNVM++: Cross-Layer Modeling and Optimization Framework of Nonvolatile Memories for Deep Learning
复制标题

DOI:
10.1109/tcad.2021.3127148
复制
发表时间:
2020-12
影响因子:
2.9
通讯作者:
A. Inci;Mehmet Meric Isgenc;Diana Marculescu
A. Inci;Mehmet Meric Isgenc;Diana Marculescu
中科院分区:
计算机科学3区
文献类型:
--
作者:
A. Inci;Mehmet Meric Isgenc;Diana Marculescu

文献摘要

被引文献

相似文献

非易失性存储器(NVM)技术,例如自旋转移扭矩磁随机访问记忆(STT-MRAM)和自旋轨道扭矩磁随机访问记忆(SOT-MRAM),与传统的SRAM相比,由于其非挥发性,具有显着优势细胞密度和可伸缩性特征。虽然先前的工作已经调查了NVM对通用应用的几种架构含义,但在这项工作中,我们提出了DEEPNVM ++,这是一个表征,模型和分析基于NVM的CACHE的框架,用于GPU架构中的GPU架构(DL)应用程序(DL)应用程序。 - 特定的电路级模型和各种DL工作负载的实际内存行为。我们介绍了依赖于常规的SRAM和新兴STT-MRAM和SOT-MRAM Technologies的系统的系统的ISO容量和ISO区域性能和能量分析。在ISO容量案例中,STT-MRAM和SOT-MRAM提供高达$ 3.8 \ times $和$ 4.7 \ times $ energy-delay产品(EDP)降低$ 2.4 \ times $和$ 2.8 \ times $ $ $ 2.8 \ times $ yable $ y \ times $ yabt , 分别。根据ISO-AREA假设,STT-MRAM和SOT-MRAM提供高达$ 2 \ times $和$ 2.3 \ times $ $ edp减少,并与SRAM相比,分别可容纳$ 2.3 \ times $和$ 3.3 \ times $ CACHE的容量。我们还执行可伸缩性分析,并表明与大型缓存能力相比,STT-MRAM和SOT-MRAM与SRAM相比实现了EDP的降低。我们的全面跨层框架在STT-/SOT-MRAM技术上进行了证明,可用于DL应用中GPU中最后一级caches的任何NVM技术的表征,建模和分析。
Nonvolatile memory (NVM) technologies, such as spin-transfer torque magnetic random access memory (STT-MRAM) and spin-orbit torque magnetic random access memory (SOT-MRAM), have significant advantages compared to conventional SRAM due to their nonvolatility, higher cell density, and scalability features. While previous work has investigated several architectural implications of NVM for generic applications, in this work, we present DeepNVM ++, a framework to characterize, model, and analyze NVM-based caches in GPU architectures for deep learning (DL) applications by combining technology-specific circuit-level models and the actual memory behavior of various DL workloads. We present both iso-capacity and iso-area performance and energy analysis for systems whose last-level caches rely on conventional SRAM and emerging STT-MRAM and SOT-MRAM technologies. In the iso-capacity case, STT-MRAM and SOT-MRAM provide up to $3.8 \times $ and $4.7 \times $ energy-delay product (EDP) reduction and $2.4 \times $ and $2.8 \times $ area reduction compared to conventional SRAM, respectively. Under iso-area assumptions, STT-MRAM and SOT-MRAM provide up to $2 \times $ and $2.3 \times $ EDP reduction and accommodate $2.3 \times $ and $3.3 \times $ cache capacity when compared to SRAM, respectively. We also perform a scalability analysis and show that STT-MRAM and SOT-MRAM achieve orders of magnitude EDP reduction when compared to SRAM for large cache capacities. Our comprehensive cross-layer framework is demonstrated on STT-/SOT-MRAM technologies and can be used for the characterization, modeling, and analysis of any NVM technology for last-level caches in GPUs for DL applications.