Compiler support for selective page migration in NUMA architectures

Compiler support for selective page migration in NUMA architectures
复制标题

编译器支持 NUMA 架构中的选择性页面迁移

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Fernando Magno Quintão Pereira
Fernando Magno Quintão Pereira
中科院分区:
--
文献类型:
--
作者:
G. Piccoli;Henrique Nazaré;R. E. Rodrigues;C. Pousa;E. Borin;Fernando Magno Quintão Pereira

文献摘要

被引文献

相似文献

当前的高性能多层处理器为用户提供了非统一的内存访问模型(NUMA) ,我们提出了编译器分析和代码生成方法,以支持轻巧的运行时系统,该系统动态迁移内存页面以改善我们的技术局部性。能够识别最有前途的页面,我们可以推断数组的大小,以及这些估算中每个内存访问指令的重复使用量。设计。我们构建动态检查的模板,只能在运行时填充它们静态分析对程序中的变量数量是二次的,并且在实践中,动态检查是O(1)。操作系统的内核。我们已经在几种平行算法上应用了我们的技术,这些算法完全忽略了不对称的记忆拓扑,并且与静态启发式方法相比,它观察到了高达4倍的加速度。我们将我们的方法与MINAS进行了比较,MINA是支持NUMA Aware数据分配的中间件,并表明我们可以在某些情况下最多优于50%。
Current high-performance multicore processors provide users with a non-uniform memory access model (NUMA). These systems perform better when threads access data on memory banks next to the core where they run. However, ensuring data locality is difficult. In this paper, we propose compiler analyses and code generation methods to support a lightweight runtime system that dynamically migrates memory pages to improve data locality. Our technique combines static and dynamic analyses and is capable of identifying the most promising pages to migrate. Statically, we infer the size of arrays, plus the amount of reuse of each memory access instruction in a program. These estimates rely on a simple, yet accurate, trip count predictor of our own design. This knowledge let's us build templates of dynamic checks, to be filled with values known only at runtime. These checks determine when it is profitable to migrate data closer to the processors where this data is used. Our static analyses are quadratic on the number of variables in a program, and the dynamic checks are O(1) in practice. Our technique does not require any form of user intervention, neither the support of a third-party middleware, nor modifications in the operating system's kernel. We have applied our technique on several parallel algorithms, which are completely oblivious to the asymmetric memory topology, and have observed speedups of up to 4×, compared to static heuristics. We compare our approach against Minas, a middleware that supports NUMA-aware data allocation, and show that we can outperform it by up to 50% in some cases.