Exploiting Page Table Locality for Agile TLB Prefetching

Exploiting Page Table Locality for Agile TLB Prefetching
复制标题

DOI:
10.1109/isca52012.2021.00016
复制
发表时间:
2021-06
期刊:
2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Georgios Vavouliotis;Lluc Alvarez;Vasileios Karakostas;K. Nikas;N. Koziris;Daniel A. Jiménez;Marc Casas
Georgios Vavouliotis;Lluc Alvarez;Vasileios Karakostas;K. Nikas;N. Koziris;Daniel A. Jiménez;Marc Casas
中科院分区:
其他
文献类型:
--
作者:
Georgios Vavouliotis;Lluc Alvarez;Vasileios Karakostas;K. Nikas;N. Koziris;Daniel A. Jiménez;Marc Casas

文献摘要

被引文献

相似文献

由于获取相应地址翻译所需的页遍历,频繁的翻译Lookaside Buffer (TLB)丢失会导致高性能和能源成本。在需求TLB访问之前预取页表项(pte)可以缓解地址转换性能瓶颈,但每次预取都需要遍历页表,从而触发对内存层次结构的额外访问。因此,TLB预取是一种代价高昂的技术,当预取不准确时可能会降低性能。本文利用页表最后一级的局部性,通过“免费”获取缓存行相邻的pte来降低TLB预取的成本,提高TLB预取的有效性。我们提出了基于采样的自由TLB预取(SBFP),这是一种动态方案,它预测这些“自由”pte的有用性,并只预取最有可能防止TLB丢失的pte。我们证明了将SBFP与新颖的、最先进的TLB预取器相结合,可以显著提高脱靶覆盖率,并减少由于页游动而导致的大部分内存访问。此外,我们提出了Agile TLB预取器(ATP),这是一种专门设计用于最大化SBFP益处的新型复合TLB预取器。ATP有效地结合了三种低成本的TLB预取器,并对那些不受益于它的执行阶段禁用TLB预取。与最先进的TLB预取器只将模式与一个特征(例如,步幅,PC,距离)相关联不同,ATP将模式与多个特征相关联,并在每次TLB丢失时动态启用最合适的TLB预取器。为了缓解地址转换性能瓶颈,我们提出了一个结合ATP和SBFP的统一解决方案。在Qualcomm提供的一组广泛的工业工作负载中,ATP与SBFP相结合将几何加速提高了16.2%,并且平均消除了由于页面行走而导致的37%的内存引用。考虑到SPEC CPU 2006和SPEC CPU 2017基准测试套件,带有SBFP的ATP几何加速提高了11.1%,消除了26%的页走内存引用。应用于大数据工作负载(GAP套件,XSBench), ATP与SBFP产生11.8%的几何加速,同时减少了5%的页面行走内存引用。对于每个基准测试套件,使用最先进的TLB预取器,对于Qualcomm、SPEC和GAP+XSBench工作负载,ATP与SBFP分别实现了8.7%、3.4%和4.2%的速度提升。
Frequent Translation Lookaside Buffer (TLB) misses incur high performance and energy costs due to page walks required for fetching the corresponding address translations. Prefetching page table entries (PTEs) ahead of demand TLB accesses can mitigate the address translation performance bottleneck, but each prefetch requires traversing the page table, triggering additional accesses to the memory hierarchy. Therefore, TLB prefetching is a costly technique that may undermine performance when the prefetches are not accurate.In this paper we exploit the locality in the last level of the page table to reduce the cost and enhance the effectiveness of TLB prefetching by fetching cache-line adjacent PTEs "for free". We propose Sampling-Based Free TLB Prefetching (SBFP), a dynamic scheme that predicts the usefulness of these "free" PTEs and prefetches only the ones most likely to prevent TLB misses. We demonstrate that combining SBFP with novel and state-of-the-art TLB prefetchers significantly improves miss coverage and reduces most memory accesses due to page walks.Moreover, we propose Agile TLB Prefetcher (ATP), a novel composite TLB prefetcher particularly designed to maximize the benefits of SBFP. ATP efficiently combines three low-cost TLB prefetchers and disables TLB prefetching for those execution phases that do not benefit from it. Unlike state-of-the-art TLB prefetchers that correlate patterns with only one feature (e.g., strides, PC, distances), ATP correlates patterns with multiple features and dynamically enables the most appropriate TLB prefetcher per TLB miss.To alleviate the address translation performance bottleneck, we propose a unified solution that combines ATP and SBFP. Across an extensive set of industrial workloads provided by Qualcomm, ATP coupled with SBFP improves geometric speedup by 16.2%, and eliminates on average 37% of the memory references due to page walks. Considering the SPEC CPU 2006 and SPEC CPU 2017 benchmark suites, ATP with SBFP increases geometric speedup by 11.1%, and eliminates page walk memory references by 26%. Applied to big data workloads (GAP suite, XSBench), ATP with SBFP yields a geometric speedup of 11.8% while reducing page walk memory references by 5%. Over the best state-of-the-art TLB prefetcher for each benchmark suite, ATP with SBFP achieves speedups of 8.7%, 3.4%, and 4.2% for the Qualcomm, SPEC, and GAP+XSBench workloads, respectively.