HALO: A Hierarchical Memory Access Locality Modeling Technique For Memory System Explorations

HALO: A Hierarchical Memory Access Locality Modeling Technique For Memory System Explorations
复制标题

DOI:
10.1145/3205289.3205323
复制
发表时间:
2018-06
期刊:
Proceedings of the 2018 International Conference on Supercomputing
影响因子:
--
通讯作者:
Reena Panda;L. John
Reena Panda;L. John
中科院分区:
其他
文献类型:
--
作者:
Reena Panda;L. John

文献摘要

被引文献

相似文献

由于其数据密集型性质,复杂的访问模式,较大的占地面积等,应用程序的日益复杂性对内存系统设计提出了新的挑战。全系统模拟器的缓慢性质,模拟器的挑战,要运行许多新兴工作量的深层软件堆栈,软件的性质等。对未来内存层次结构的快速准确的微体系探索构成挑战。缓解此问题的一种技术是创建访问流的时空模型,并使用它们来探索内存系统权衡。但是,现有的内存流模型只有使用全局阶段过渡模型的时间局部性行为或模型时空位置对它们进行建模,从而导致高存储/元数据开销。在本文中,我们提出了Halo,这是一种层次内存访问局部建模技术,该技术通过将全局内存引用隔离到本地化流中,并进一步放大到每个本地流中,从而识别模式。 Halo还建模了局部流访问之间的交织程度,可利用粗粒的再利用区域。我们评估了Halo在使用20K以上不同的内存系统配置中复制原始应用程序性能方面的有效性,并表明Halo在复制启用Prefetcher L1&L2 Caches,TLB和DRAM,TLB和DRAM,TLB和DRAM,TLB,TLB和DRAM的绩效方面的精度超过98.3%,95.6%,99.3%和96%的精度。 Halo的表现优于最先进的内存克隆方案,WEST和STM,同时使用的元数据储存量比STM少约39倍。
Growing complexity of applications pose new challenges to memory system design due to their data intensive nature, complex access patterns, larger footprints, etc. The slow nature of full-system simulators, challenges of simulators to run deep software stacks of many emerging workloads, proprietary nature of software, etc. pose challenges to fast and accurate microarchitectural explorations of future memory hierarchies. One technique to mitigate this problem is to create spatio-temporal models of access streams and use them to explore memory system tradeoffs. However, existing memory stream models have weaknesses such as they only model temporal locality behavior or model spatio-temporal locality using global stride transitions, resulting in high storage/metadata overhead. In this paper, we propose HALO, a Hierarchical memory Access LOcality modeling technique that identifies patterns by isolating global memory references into localized streams and further zooming into each local stream capturing multi-granularity spatial locality patterns. HALO also models the interleaving degree between localized stream accesses leveraging coarse-grained reuse locality. We evaluate HALO's effectiveness in replicating original application performance using over 20K different memory system configurations and show that HALO achieves over 98.3%, 95.6%, 99.3% and 96% accuracy in replicating performance of prefetcher-enabled L1 & L2 caches, TLB and DRAM respectively. HALO outperforms the state-of-the-art memory cloning schemes, WEST and STM, while using ~39X less metadata storage than STM.