Optimizing the TLB Shootdown Algorithm with Page Access Tracking

Optimizing the TLB Shootdown Algorithm with Page Access Tracking
复制标题

通过页面访问跟踪优化 TLB 击落算法

DOI:
--
复制
发表时间:
2017
期刊:
USENIX Annual Technical Conference
影响因子:
--
通讯作者:
Nadav Amit
Nadav Amit
中科院分区:
--
文献类型:
--
作者:
Nadav Amit

文献摘要

被引文献

相似文献

操作系统的任务是维护每核TLB的一致性,需要昂贵的同步操作,特别是使陈旧的映射无效。随着内核数量的增加,TLB同步的开销也会增加,并阻碍可扩展性,而现有的软件优化,试图减轻问题(如缓存)是缺乏的。我们通过修改TLB同步子系统来解决这个问题。我们介绍了几种技术,检测的情况下,即将被无效的映射缓存只有一个TLB或根本没有缓存,使我们能够完全避免同步的成本。与现有的优化相比,我们的方法利用硬件页面访问跟踪。我们在Linux中实现了我们的技术,发现它们平均减少了高达98%的TLB无效次数,从而提高了高达78%的性能。评估表明,虽然我们的技术可能会引入高达9%的内存映射时,永远不会删除的开销,这些开销可以避免简单的硬件增强。
The operating system is tasked with maintaining the coherency of per-core TLBs, necessitating costly synchronization operations, notably to invalidate stale mappings. As core-counts increase, the overhead of TLB synchronization likewise increases and hinders scalability, whereas existing software optimizations that attempt to alleviate the problem (like batching) are lacking. We address this problem by revising the TLB synchronization subsystem. We introduce several techniques that detect cases whereby soon-to-be invalidated mappings are cached by only one TLB or not cached at all, allowing us to entirely avoid the cost of synchronization. In contrast to existing optimizations, our approach leverages hardware page access tracking. We implement our techniques in Linux and find that they reduce the number of TLB invalidations by up to 98% on average and thus improve performance by up to 78%. Evaluations show that while our techniques may introduce overheads of up to 9% when memory mappings are never removed, these overheads can be avoided by simple hardware enhancements.