Performance and Energy-Efficient Design of STT-RAM Last-Level Cache

Performance and Energy-Efficient Design of STT-RAM Last-Level Cache
复制标题

DOI:
10.1109/tvlsi.2018.2804938
复制
发表时间:
2018-03
影响因子:
2.8
通讯作者:
F. Hameed;A. Khan;J. Castrillón
F. Hameed;A. Khan;J. Castrillón
中科院分区:
工程技术2区
文献类型:
--
作者:
F. Hameed;A. Khan;J. Castrillón

文献摘要

被引文献

相似文献

最近的研究提出了一种芯片堆叠式末级高速缓存(LLC)来克服存储墙。最近,自旋转移扭矩随机存取存储器(STT-RAM)缓存受到了关注,因为与DRAM缓存相比,它们提供了更高的能效。然而,最近提出的STT-RAM高速缓存体系结构通过将不需要的高速缓存线(CL)取到行缓冲器(RB)中来不必要地消耗能量。在本文中,我们提出了一种针对STT-RAM的选择性读取策略,该策略将那些可能被重用的CLS取出到RB中。此外,我们还提出了一种减少STT-RAM回写次数的标签更新策略。这减少了读/写次数,从而降低了能耗。为了减少选择性读策略的延迟代价,我们提出了以下性能优化:1)减少STT-RAM访问延迟的RB标签旁路策略;2)存储可能在不久的将来使用的CLS的LLC数据缓存;3)同时降低LLC访问延迟和错失率的地址组织方案;以及4)提高访问并行性的标签到列映射策略。为了进行评估,我们在Zesto模拟器中实现了我们提出的体系结构,并在一个八核系统上运行了不同的SPEC2006基准测试组合。我们与最近提出的对子阵列并行支持的STT-RAM LLC进行了比较,结果表明,我们的协同策略使LLC的平均动态能耗降低了75%,系统性能提高了6.5%。与目前最先进的子阵列并行DRAM LLC相比,LLC的动态能耗降低了82%,系统性能提高了6.8%。
Recent research has proposed having a die-stacked last-level cache (LLC) to overcome the memory wall. Lately, spin-transfer-torque random access memory (STT-RAM) caches have received attention, since they provide improved energy efficiency compared with DRAM caches. However, recently proposed STT-RAM cache architectures unnecessarily dissipate energy by fetching unneeded cache lines (CLs) into the row buffer (RB). In this paper, we propose a selective read policy for the STT-RAM which fetches those CLs into the RB that are likely to be reused. In addition, we propose a tags-update policy that reduces the number of STT-RAM writebacks. This reduces the number of reads/writes and thereby decreases the energy consumption. To reduce the latency penalty of our selective read policy, we propose the following performance optimizations: 1) an RB tags-bypass policy that reduces STT-RAM access latency; 2) an LLC data cache that stores the CLs that are likely to be used in the near future; 3) an address organization scheme that simultaneously reduces LLC access latency and miss rate; and 4) a tags-to-column mapping policy that improves access parallelism. For evaluation, we implement our proposed architecture in the Zesto simulator and run different combinations of SPEC2006 benchmarks on an eight-core system. We compare our approach with a recently proposed STT-RAM LLC with subarray parallelism support and show that our synergistic policies reduce the average LLC dynamic energy consumption by 75% and improve the system performance by 6.5%. Compared with the state-of-the-art DRAM LLC with subarray parallelism, our architecture reduces the LLC dynamic energy consumption by 82% and improves system performance by 6.8%.