The Impact of Page Size and Microarchitecture on Instruction Address Translation Overhead

The Impact of Page Size and Microarchitecture on Instruction Address Translation Overhead
复制标题

DOI:
10.1145/3600089
复制
发表时间:
2023-05
影响因子:
1.6
通讯作者:
Yufeng Zhou;A. Cox;S. Dwarkadas;Xiaowan Dong
Yufeng Zhou;A. Cox;S. Dwarkadas;Xiaowan Dong
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yufeng Zhou;A. Cox;S. Dwarkadas;Xiaowan Dong

文献摘要

被引文献

相似文献

随着应用程序处理的数据量的增加,数据地址转换开销受到了相当大的关注,导致了较大页面大小(超页)和多级转换后备缓冲器(TLB)的广泛使用。然而,对指令地址翻译及其与TLB和流水线结构的关系的关注要少得多。在之前的工作中,我们量化了使用代码超页对各种广泛使用的应用程序的影响,从编译器到Web用户界面框架,以及为可执行文件和共享库共享页面表页的影响。在本文中,我们首先揭示Intel Skylake和AMD Zen+之间的微体系结构差异,特别是它们不同的TLB组织对指令地址转换开销的影响,从而增强这些结果。这一分析为影响指令地址转换成本的微体系结构设计决策提供了一些关键的见解。首先,在使用代码超页时,指令和数据映射同时竞争同一结构中的空间的较低级别(2级)TLB可以实现更好的整体性能和利用率。代码超页不仅减少了指令地址转换开销,还间接减少了数据地址转换开销。事实上,对于少数应用程序,仅使用几个代码超页对整体性能的影响比使用数量大得多的数据超页要大得多。其次,具有用于不同页大小的不同结构的1级(L1)TLB可能需要仔细地调整代码的超页提升策略,并且相应地需要对2级TLB的次最佳利用。具体地说,当L1超页结构的大小较小时增加超页的数量可能会导致某些应用程序的L1 TLB未命中更多。此外,在一些微体系结构上,这些未命中的成本可能是高度可变的,因为替换被延迟,直到受害者条目映射的所有正在运行的指令都退役。因此,更多的超级页面促销可能会导致性能倒退。最后,我们的发现也证明了操作系统对包含可执行文件和共享库的普通文件上的超页的一流支持,以及对代码的更积极的超页策略。
As the volume of data processed by applications has increased, considerable attention has been paid to data address translation overheads, leading to the widespread use of larger page sizes (“superpages”) and multi-level translation lookaside buffers (TLBs). However, far less attention has been paid to instruction address translation and its relation to TLB and pipeline structure. In prior work, we quantified the impact of using code superpages on a variety of widely used applications, ranging from compilers to web user-interface frameworks, and the impact of sharing page table pages for executables and shared libraries. Within this article, we augment those results by first uncovering the effects that microarchitectural differences between Intel Skylake and AMD Zen+, particularly their different TLB organizations, have on instruction address translation overhead. This analysis provides some key insights into the microarchitectural design decisions that impact the cost of instruction address translation. First, a lower-level (level 2) TLB that has both instruction and data mappings competing for space within the same structure allows better overall performance and utilization when using code superpages. Code superpages not only reduce instruction address translation overhead but also indirectly reduce data address translation overhead. In fact, for a few applications, the use of just a few code superpages has a larger impact on overall performance than the use of a much larger number of data superpages. Second, a level 1 (L1) TLB with separate structures for different page sizes may require careful tuning of the superpage promotion policy for code, and a correspondingly suboptimal utilization of the level 2 TLB. In particular, increasing the number of superpages when the size of the L1 superpage structure is small may result in more L1 TLB misses for some applications. Moreover, on some microarchitectures, the cost of these misses can be highly variable, because replacement is delayed until all of the in-flight instructions mapped by the victim entry are retired. Hence, more superpage promotions can result in a performance regression. Finally, our findings also make a case for first-class OS support for superpages on ordinary files containing executables and shared libraries, as well as a more aggressive superpage policy for code.