Mosaic Pages: Big TLB Reach with Small Pages

Mosaic Pages: Big TLB Reach with Small Pages
复制标题

马赛克页面:小页面实现大 TLB 覆盖范围

DOI:
10.1145/3582016.3582021
复制
发表时间:
2023
期刊:
Proc.\ 28th {ACM} International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Bhattacharjee, Abhishek
Bhattacharjee, Abhishek
中科院分区:
--
文献类型:
--
作者:
Gosakan, Krishnan;Han, Jaehyun;Kuszmaul, William;Mubarek, Ibrahim N.;Mukherjee, Nirjhar;Sriram, Karthik;Tagliavini, Guido;West, Evan;Bender, Michael A.;Bhattacharjee, Abhishek

文献摘要

参考文献

被引文献

相似文献

TLB越来越成为大数据应用的瓶颈。在大多数设计中,TLB条目的数量受到延迟要求的高度限制,并且比应用程序的工作集增长得慢得多。这个问题的许多解决方案,如巨大的页面,穿孔的页面,或TLB合并,依赖于物理连续性的性能增益,但碎片整理内存的成本可以很容易地抵消这些增益。本文介绍了镶嵌页,它通过将多个离散的翻译压缩到一个TLB条目中来增加TLB范围。 Mosaic利用虚拟邻接来实现局部性,但不使用物理邻接。Mosaic依赖于哈希理论的最新进展来约束内存映射,以便在不降低内存利用率或增加交换的情况下实现这种物理地址压缩。本文介绍了一个完整的系统原型马赛克,在gem 5和修改后的Linux。在模拟中,与传统设计相比,Mosaic将多个工作负载中的TLB未命中减少了6- 81%。我们的研究结果表明,Mosaic对内存映射的限制不会损害性能,在我们的实验中,在内存达到98%之前,我们从未看到冲突-在这一点上,传统的设计也可能会交换。一旦内存过度使用,Mosaic在大多数情况下交换的页面比Linux少。最后,我们提出了时间和面积分析的verilog实现所需的关键路径上的TLB的散列函数,并表明,在商业28纳米CMOS工艺,电路运行在4 GHz的最大频率,表明马赛克TLB是不太可能影响时钟频率。
The TLB is increasingly a bottleneck for big data applications. In most designs, the number of TLB entries are highly constrained by latency requirements, and growing much more slowly than the working sets of applications. Many solutions to this problem, such as huge pages, perforated pages, or TLB coalescing, rely on physical contiguity for performance gains, yet the cost of defragmenting memory can easily nullify these gains. This paper introduces mosaic pages, which increase TLB reach by compressing multiple, discrete translations into one TLB entry. Mosaic leverages virtual contiguity for locality, but does not use physical contiguity. Mosaic relies on recent advances in hashing theory to constrain memory mappings, in order to realize this physical address compression without reducing memory utilization or increasing swapping. This paper presents a full-system prototype of Mosaic, in gem5 and modified Linux. In simulation and with comparable hardware to a traditional design, mosaic reduces TLB misses in several workloads by 6-81%. Our results show that Mosaic’s constraints on memory mappings do not harm performance, we never see conflicts before memory is 98% full in our experiments — at which point, a traditional design would also likely swap. Once memory is over-committed, Mosaic swaps fewer pages than Linux in most cases. Finally, we present timing and area analysis for a verilog implementation of the hashing function required on the critical path for the TLB, and show that on a commercial 28nm CMOS process; the circuit runs at a maximum frequency of 4 GHz, indicating that a mosaic TLB is unlikely to affect clock frequency.
Linux / ia 64 项目:内核设计和状态更新
DOI: --
发表时间: 2000
期刊:
影响因子: --
作者:
S. Eranian;D. Mosberger
通讯作者: D. Mosberger
通过机会虚拟缓存减少内存引用能量
DOI: 10.1145/2366231.2337194
发表时间: 2012
期刊: 2012 39th Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者:
Arkaprava Basu;M. Hill;M. Swift
通讯作者: M. Swift
多处理器虚拟地址缓存的一致性
DOI: --
发表时间: 1987
期刊: International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者:
J. Goodman
通讯作者: J. Goodman
重新审视虚拟 L1 缓存:使用动态同义词重新映射的实用设计
DOI: --
发表时间: 2016
期刊: International Symposium on High-Performance Computer Architecture
影响因子: --
作者:
Hongil Yoon;G. Sohi
通讯作者: G. Sohi
DOI: 10.1145/3409964.3461814
发表时间: 2021-07
期刊: Proceedings of the 33rd ACM Symposium on Parallelism in Algorithms and Architectures
影响因子: --
作者:
M. A. Bender;A. Bhattacharjee;Alex Conway;Martín Farach-Colton;Rob Johnson;Sudarsun Kannan;William Kuszmaul;Nirjhar Mukherjee;Donald E. Porter;Guido Tagliavini;Janet Vorobyeva;Evan West
通讯作者: M. A. Bender;A. Bhattacharjee;Alex Conway;Martín Farach-Colton;Rob Johnson;Sudarsun Kannan;William Kuszmaul;Nirjhar Mukherjee;Donald E. Porter;Guido Tagliavini;Janet Vorobyeva;Evan West