A Comparative Analysis on the Impact of Bank Contention in STT-MRAM and SRAM Based LLCs

A Comparative Analysis on the Impact of Bank Contention in STT-MRAM and SRAM Based LLCs
复制标题

基于 STT-MRAM 和 SRAM 的有限责任公司银行竞争影响的比较分析

DOI:
10.1109/iccd46524.2019.00039
复制
发表时间:
2019
期刊:
2019 IEEE 37th International Conference on Computer Design (ICCD)
影响因子:
--
通讯作者:
José Ignacio Gómez Pérez
José Ignacio Gómez Pérez
中科院分区:
--
文献类型:
--
作者:
Timon Evenblij;C. Tenllado;M. Perumkunnil;F. Catthoor;S. Sakhare;P. Debacker;G. Kar;A. Furnémont;Nicolas Bueno;José Ignacio Gómez Pérez

文献摘要

参考文献

被引文献

相似文献

自旋转移扭矩磁 RAM (STT-MRAM) 由于其高密度、低泄漏和非易失性而被广泛认为是末级缓存 (LLC) 的有前途的替代品。然而,对 STT-MRAM 的写入是能源密集型的并且具有很高的延迟。虽然写入期间的高动态能耗可以通过低静态能耗来补偿,但高延迟会导致性能下降。这项工作表明,与基于 SRAM 的 LLC 相比,STT-MRAM 的性能下降主要是由于在写入存储体时尝试满足读取请求时的存储体争用造成的。我们全面探讨了高速缓存组和高速缓存争用对具有有序内核或无序内核的移动多核系统 LLC 中的能量和性能的影响。分析的细节是通过基于 28nm SRAM 行业编译器和内部开发的 STT-MRAM 编译器的高精度缓存模型实现的,该编译器生成完整的 STT-MRAM 宏设计,具有经过硅验证的 MTJ 堆栈和 28nm 节点上的完整寄生提取。我们的结果表明,STT-MRAM 缓存和 SRAM 缓存之间的能源性能最佳存储配置存在明显差异。与 SRAM 缓存设计相比,这些低争用 STT-MRAM 缓存设计具有最佳存储体数量,可节省至少 60% 的缓存能量,同时系统性能损失最多个位数百分比。这表明在 LLC 中使用 STT-MRAM 替代 SRAM 的潜力越来越大。
Spin Transfer Torque Magnetic RAM (STT-MRAM) is being extensively considered as a promising replacement for Last Level Caches (LLC), due to its high density, low leakage and non-volatility. However, writes to STT-MRAM are energy intensive and have a high latency. While the high dynamic energy consumption during writes can be compensated by the low static energy consumption, the high latency results in performance degradation. This work shows that in contrast to SRAM-based LLCs, the performance degradation for STT-MRAM is primarily due to bank contention, when trying to satisfy a read request while the bank is being written. We holistically explore the effects of cache banking and cache contention on energy and performance in the LLC of mobile multicore systems, with in-order cores or with out-of-order cores. The detail of the analysis is enabled by highly accurate cache models, based on a 28nm SRAM industry compiler, and an in-house developed STT-MRAM compiler, which generates full STT-MRAM macro designs with silicon-validated MTJ stack and complete parasitic extraction at the 28nm node. Our results show that there is a clear difference in the energy-performance optimal banking configuration between STT-MRAM caches and SRAM caches. These low contention STT-MRAM cache designs with the optimal number of banks save at least 60% cache energy while losing at most single digit percentages in system performance compared to SRAM cache designs. This show an increased potential of using STT-MRAM as a replacement for SRAM in an LLC.
DOI: 10.1109/tvlsi.2018.2804938
发表时间: 2018-03
影响因子: 2.8
作者:
F. Hameed;A. Khan;J. Castrillón
通讯作者: F. Hameed;A. Khan;J. Castrillón