Toward multi-programmed workloads with different memory footprints: a self-adaptive last level cache scheduling scheme

Toward multi-programmed workloads with different memory footprints: a self-adaptive last level cache scheduling scheme
复制标题

DOI:
10.1007/s11432-016-0408-1
复制
发表时间:
2017-07
期刊:
Science China Information Sciences
影响因子:
--
通讯作者:
Jingyu Zhang;M. Guo;Chentao Wu;Yuanyi Chen
Jingyu Zhang;M. Guo;Chentao Wu;Yuanyi Chen
中科院分区:
其他
文献类型:
--
作者:
Jingyu Zhang;M. Guo;Chentao Wu;Yuanyi Chen

文献摘要

被引文献

相似文献

随着3D堆叠技术的出现,动态随机存取存储器(DRAM)可以堆叠在芯片上来构建DRAM末级缓存(LLC)。与静态随机存取存储器(SRAM)相比,DRAM更大但速度更慢。在现有的研究论文中,从SRAM结构的改进到缓存标记和数据访问的优化,很多工作都致力于将SRAM和堆叠DRAM结合起来来提高工作负载性能。相反,很少有人关注为具有不同内存占用的多程序工作负载设计LLC调度方案。受此启发,我们提出了一种自适应的LLC调度方案,使我们能够高效地利用SRAM和3D堆叠的DRAM,从而获得更好的工作负载性能。该调度方案采用了(1)评估单元,用于探测和评估程序执行过程中的缓存信息;(2)实现单元,用于自适应地选择SRAM或DRAM。为了使调度方案能够正常工作,我们制定了数据迁移策略。我们进行了大量的实验来评估我们所提出的方案的性能。实验结果表明,与现有方法相比,该方法可以将多程序负载性能提高30%以上。
With the emerging of 3D-stacking technology, the dynamic random-access memory (DRAM) can be stacked on chips to architect the DRAM last level cache (LLC). Compared with static randomaccess memory (SRAM), DRAM is larger but slower. In the existing research papers, a lot of work has been devoted to improving the workload performance using SRAM and stacked DRAM together, ranging from SRAM structure improvement, to optimizing cache tag and data access. Instead, little attention has been paid to designing an LLC scheduling scheme for multi-programmed workloads with different memory footprints. Motivated by this, we propose a self-adaptive LLC scheduling scheme, which allows us to utilize SRAM and 3D-stacked DRAM efficiently, achieving better workload performance. This scheduling scheme employs (1) an evaluation unit, which is used to probe and evaluate the cache information during the process of programs being executed; and (2) an implementation unit, which is used to self-adaptively choose SRAM or DRAM. To make the scheduling scheme work correctly, we develop a data migration policy. We conduct extensive experiments to evaluate the performance of our proposed scheme. Experimental results show that our method can improve the multi-programmed workload performance by up to 30% compared with the state-of-the-art methods.