Memory management in NUMA multicore systems: trapped between cache contention and interconnect overhead

Memory management in NUMA multicore systems: trapped between cache contention and interconnect overhead
复制标题

DOI:
10.1145/1993478.1993481
复制
发表时间:
2011-06
影响因子:
3.2
通讯作者:
Z. Majó;T. Gross
Z. Majó;T. Gross
中科院分区:
地球科学3区
文献类型:
--
作者:
Z. Majó;T. Gross

文献摘要

被引文献

相似文献

基于多核处理器的多处理器通常包括非均匀存储器架构(NUMA);即使是当前的8核双处理器系统也表现出非均匀的存储器访问时间。由于处理器的内核共享一个公共缓存,因此必须重新考虑内存管理和进程映射的问题。我们发现,优化数据局部性可以抵消缓存争用避免的好处,反之亦然。因此,系统软件必须同时考虑数据局部性和缓存争用以实现良好的性能,并且内存管理不能与进程调度分离。我们提出了一个商业上可用的NUMA多核架构,英特尔Nehalem的详细分析。我们描述了两个调度算法:最大本地,优化最大的数据本地性,其扩展,N-MASS,减少数据本地性,以避免缓存争用造成的性能下降。N-MASS经过微调,可支持NUMA多核上的内存管理,与当前Linux实现中的默认设置相比,性能最多可提高32%,平均提高7%。
Multiprocessors based on processors with multiple cores usually include a non-uniform memory architecture (NUMA); even current 2-processor systems with 8 cores exhibit non-uniform memory access times. As the cores of a processor share a common cache, the issues of memory management and process mapping must be revisited. We find that optimizing only for data locality can counteract the benefits of cache contention avoidance and vice versa. Therefore, system software must take both data locality and cache contention into account to achieve good performance, and memory management cannot be decoupled from process scheduling. We present a detailed analysis of a commercially available NUMA-multicore architecture, the Intel Nehalem. We describe two scheduling algorithms: maximum-local, which optimizes for maximum data locality, and its extension, N-MASS, which reduces data locality to avoid the performance degradation caused by cache contention. N-MASS is fine-tuned to support memory management on NUMA-multicores and improves performance up to 32%, and 7% on average, over the default setup in current Linux implementations.