Efficient Metadata Management for Irregular Data Prefetching

Efficient Metadata Management for Irregular Data Prefetching
复制标题

DOI:
10.1145/3307650.3322225
复制
发表时间:
2019-06
期刊:
2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Hao Wu;Krishnendra Nathella;Dam Sunwoo;Akanksha Jain;Calvin Lin
Hao Wu;Krishnendra Nathella;Dam Sunwoo;Akanksha Jain;Calvin Lin
中科院分区:
其他
文献类型:
--
作者:
Hao Wu;Krishnendra Nathella;Dam Sunwoo;Akanksha Jain;Calvin Lin

文献摘要

相似文献

暂时的预脱水有可能预取任意的内存访问模式,但是它们需要大量的元数据,这些元数据通常必须存储在DRAM中。 2013年,不规则的流缓冲液(ISB)展示了如何通过将其内容与TLB同步,可以在芯片上缓存并隐式管理。本文揭示了该方法的效率低下,并提出了一种新的元数据管理方案,该方案使用简单的元数据预摘要来喂食元数据缓存。结果是托管的ISB(MISB),这是一种时间预摘要,在交通开销和IPC方面都可以显着前进。使用高度准确的专有模拟器进行单核工作负载,并使用Champsim模拟器进行多核工作负载,我们评估了Spec CPU 2006和CloudSuite Benchmarks Suites的程序中的MISB。我们的结果表明,对于单核工作负载,MISB提高了22.7%的绩效,而理想化的STM为10.6%,而现实的ISB则提高了4.5%。 MISB还大大降低了芯片外流量;对于Spec而言,MISB的交通开销为70%,约占STM的五分之一(342%)和ISB的六分之一(411%)。在4核多编程工作负载上,MISB提高了19.9%的绩效,而理想化的STM为7.5%。对于CloudSuite,MISB将性能提高了7.2%(理想化的STM的4.0%),同时达到了11倍的交通量(MISB的96.2%,而STMS的交通量为1082.7%)。
Temporal prefetchers have the potential to prefetch arbitrary memory access patterns, but they require large amounts of metadata that must typically be stored in DRAM. In 2013, the Irregular Stream Buffer (ISB), showed how this metadata could be cached on chip and managed implicitly by synchronizing its contents with that of the TLB. This paper reveals the inefficiency of that approach and presents a new metadata management scheme that uses a simple metadata prefetcher to feed the metadata cache. The result is the Managed ISB (MISB), a temporal prefetcher that significantly advances the state-of-the-art in terms of both traffic overhead and IPC. Using a highly accurate proprietary simulator for single-core workloads, and using the ChampSim simulator for multi-core workloads, we evaluate MISB on programs from the SPEC CPU 2006 and CloudSuite benchmarks suites. Our results show that for single-core workloads, MISB improves performance by 22.7%, compared to 10.6% for an idealized STMS and 4.5% for a realistic ISB. MISB also significantly reduces off-chip traffic; for SPEC, MISB's traffic overhead of 70% is roughly one fifth of STMS's (342%) and one sixth of ISB's (411%). On 4-core multi-programmed workloads, MISB improves performance by 19.9%, compared to 7.5% for idealized STMS. For CloudSuite, MISB improves performance by 7.2% (vs. 4.0% for idealized STMS), while achieving a traffic reduction of 11x (96.2% for MISB vs. 1082.7% for STMS).