Prefetched Address Translation

Prefetched Address Translation
复制标题

预取地址转换

DOI:
10.1145/3352460.3358294
复制
发表时间:
2019
期刊:
Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Boris Grot
Boris Grot
中科院分区:
--
文献类型:
--
作者:
Artemiy Margaritov;Dmitrii Ustiugov;Edouard Bugnion;Boris Grot

文献摘要

被引文献

相似文献

随着数据集大小的爆炸性增长和机器内存能力的增加,每个应用记忆足迹通常会进入数百个GB。如此巨大的数据集向TLB施加压力,导致经常失踪必须通过页面步行解决 - 长期长寿指针通过基于内存的Radix树的多个级别追逐基于内存的Radix树的页面表。这项工作预计数据集大小的进一步增长及其对TLB命中率的不利影响,该工作旨在加速页面步行,同时完全保留现有的虚拟内存抽象和机​​制 - 对于软件兼容性和一般性的必不可少。我们的想法是将直接索引置入给定的页面表的级别,从而需要首先从前一个级别获取指针。我们工作的一个关键贡献是表明可以通过简单地订购包含Page表中的PAGE表以匹配其映射到的虚拟内存页面的顺序来完成。这样做可以使用基础加上算盘算术直接索引到页面表。我们介绍了用预取(ASAP)的地址转换,这是一种将地址转换延迟到对内存层次结构的单一访问的新方法。在TLB错过时,ASAP将预取式供应到页面表的更深层次,从而绕过了前面的级别。这些预购与传统的页面步行同时发生,该步行是由于预摘要而导致的潜伏期减小,同时保证只能食用正确预测的条目。尽快需要对OS和微不足道的微体系支撑的最小扩展。此外,ASAP是完全保证的,不需要对现有的基于Radix树的页面表,TLB和其他软件和硬件机制进行地址翻译的修改。我们对一系列记忆密集型工作负载的评估表明,在SMT托管下,ASAP能够在本机执行中平均减少25%(最大42%),而在虚拟化下,最大45%(55%)。
With explosive growth in dataset sizes and increasing machine memory capacities, per-application memory footprints are commonly reaching into hundreds of GBs. Such huge datasets pressure the TLB, resulting in frequent misses that must be resolved through a page walk -- a long-latency pointer chase through multiple levels of the in-memory radix tree-based page table. Anticipating further growth in dataset sizes and their adverse affect on TLB hit rates, this work seeks to accelerate page walks while fully preserving existing virtual memory abstractions and mechanisms -- a must for software compatibility and generality. Our idea is to enable direct indexing into a given level of the page table, thus eliding the need to first fetch pointers from the preceding levels. A key contribution of our work is in showing that this can be done by simply ordering the pages containing the page table in physical memory to match the order of the virtual memory pages they map to. Doing so enables direct indexing into the page table using a base-plus-offset arithmetic. We introduce Address Translation with Prefetching (ASAP), a new approach for reducing the latency of address translation to a single access to the memory hierarchy. Upon a TLB miss, ASAP launches prefetches to the deeper levels of the page table, bypassing the preceding levels. These prefetches happen concurrently with a conventional page walk, which observes a latency reduction due to prefetching while guaranteeing that only correctly-predicted entries are consumed. ASAP requires minimal extensions to the OS and trivial microarchitectural support. Moreover, ASAP is fully legacy-preserving, requiring no modifications to the existing radix tree-based page table, TLBs and other software and hardware mechanisms for address translation. Our evaluation on a range of memory-intensive workloads shows that under SMT colocation, ASAP is able to reduce page walk latency by an average of 25% (42% max) in native execution, and 45% (55% max) under virtualization.