SELF: A High Performance and Bandwidth Efficient Approach to Exploiting Die-Stacked DRAM as Part of Memory

SELF: A High Performance and Bandwidth Efficient Approach to Exploiting Die-Stacked DRAM as Part of Memory
复制标题

SELF:一种利用裸片堆叠 DRAM 作为内存一部分的高性能和带宽高效方法

DOI:
10.1109/mascots.2017.23
复制
发表时间:
2017
期刊:
and Simulation of Computer and Telecommunication Systems (MASCOTS
影响因子:
--
通讯作者:
He, Xubin
He, Xubin
中科院分区:
--
文献类型:
--
作者:
Guo, Yuhua;Liu, Qing;Xiao, Weijun;Huang, Ping;Podhorszki, Norbert;Klasky, Scott;He, Xubin

文献摘要

参考文献

被引文献

相似文献

芯片堆叠式DRAM(也称为片上DRAM)提供比片外DRAM高得多的带宽和更低的延迟。这是一项很有前途的打破“记忆墙”的技术。芯片堆叠的DRAM既可以用作高速缓存(即,DRAM高速缓存),也可以用作存储器(POM)的一部分。DRAM高速缓存设计将遭受比POM设计更多的页面错误,因为DRAM高速缓存不能贡献主存储器的容量。同时,要获得高性能,需要POM系统将请求的数据交换到芯片堆叠的DRAM。现有的POM设计分为两类-基于行的设计和基于页面的设计。前者保证了低的片外带宽利用率,但由于时间局部性的限制,片上存储器的命中率很低。相比之下,基于页面的设计实现了片上存储器的高命中率,但代价是在片上和片外存储器之间移动大量数据,导致片外带宽利用率增加和系统性能显著下降。为了获得与基于页的设计相似的片上存储器的高命中率,并消除所涉及的过多的片外流量,我们提出了一种高性能和带宽高效的方法SELF。其关键思想是有选择地交换请求页面中可能根据页面足迹访问的行,而不是盲目地交换整个页面。在这样做的同时,SELF允许从片上存储器尽可能地处理传入的请求,同时避免交换未使用的行以减少存储器带宽消耗。我们评估了一个由4 GB片上DRAM和12 GB片外DRAM组成的存储系统。与总容量相同的16 GB片外DRAM的基准系统相比,SELF在每周期指令数方面的性能提高了26.9%,每次内存访问的能耗平均降低了47.9%。相比之下,与相同的基准系统相比,最先进的基于线条和基于页面的POM设计只能将性能分别提高9.5%和9.9%。
Die-stacked DRAM (a.k.a., on-chip DRAM) provides much higher bandwidth and lower latency than off-chip DRAM. It is a promising technology to break the "memory wall". Die-stacked DRAM can be used either as a cache (i.e., DRAM cache) or as a part of memory (PoM). A DRAM cache design would suffer from more page faults than a PoM design as the DRAM cache cannot contribute towards capacity of main memory. At the same time, obtaining high performance requires PoM systems to swap requested data to the die-stacked DRAM. Existing PoM designs fall into two categories – line-based and page-based. The former ensures low off-chip bandwidth utilization but suffers from a low hit ratio of on-chip memory due to limited temporal locality. In contrast, page-based designs achieve a high hit ratio of on-chip memory albeit at the cost of moving large amounts of data between on-chip and off-chip memories, leading to increased off-chip bandwidth utilization and significant system performance degradation.To achieve a similar high hit ratio of on-chip memory as page-based designs, and eliminate excessive off-chip traffic involved, we propose SELF, a high performance and bandwidth efficient approach. The key idea is to SElectively swap Lines in a requested page that are likely to be accessed according to page Footprint, instead of blindly swapping an entire page. In doing so, SELF allows incoming requests to be serviced from the on-chip memory as much as possible, while avoiding swapping unused lines to reduce memory bandwidth consumption. We evaluate a memory system which consists of 4GB on-chip DRAM and 12GB off-chip DRAM. Compared to a baseline system that has the same total capacity of 16GB off-chip DRAM, SELF improves the performance in terms of instructions per cycle by 26.9%, and reduces the energy consumption per memory access by 47.9% on average. In contrast, state-of-the-art line-based and page-based PoM designs can only improve the performance by 9.5% and 9.9%, respectively, against the same baseline system.
标签表
DOI: --
发表时间: 2015
期刊: International Symposium on High-Performance Computer Architecture
影响因子: --
作者:
Sean Franey;Mikko H. Lipasti
通讯作者: Mikko H. Lipasti
准确且复杂有效的空间模式预测
DOI: --
发表时间: 2004
期刊: International Symposium on High-Performance Computer Architecture
影响因子: --
作者:
Chi F. Chen;Se;B. Falsafi;Andreas Moshovos
通讯作者: Andreas Moshovos
用于高性能 3D DRAM 架构的历史辅助自适应粒度缓存 (HAAG$)
DOI: --
发表时间: 2015
期刊: International Conference on Supercomputing
影响因子: --
作者:
Ke Chen;Sheng Li;Jung Ho Ahn;N. Muralimanohar;Jishen Zhao;Cong Xu;O. Seongil;Yuan Xie;J. Brockman;N. Jouppi
通讯作者: N. Jouppi