A Hierarchical Approach for Load Balancing on Parallel Multi-core Systems

A Hierarchical Approach for Load Balancing on Parallel Multi-core Systems
复制标题

DOI:
10.1109/icpp.2012.9
复制
发表时间:
2012-09
期刊:
2012 41st International Conference on Parallel Processing
影响因子:
--
通讯作者:
L. Pilla;Christiane Pousa Ribeiro;Daniel Cordeiro;Chao Mei;A. Bhatele;P. Navaux;François Broquedis;Jean-François Méhaut;L. Kalé
L. Pilla;Christiane Pousa Ribeiro;Daniel Cordeiro;Chao Mei;A. Bhatele;P. Navaux;François Broquedis;Jean-François Méhaut;L. Kalé
中科院分区:
其他
文献类型:
--
作者:
L. Pilla;Christiane Pousa Ribeiro;Daniel Cordeiro;Chao Mei;A. Bhatele;P. Navaux;François Broquedis;Jean-François Méhaut;L. Kalé

文献摘要

被引文献

相似文献

具有非均匀内存访问(NUMA)的多核计算节点现在是大型并行机器组装中的一种常见架构。在这些机器上,除了网络通信成本之外,计算节点内的内存访问成本也是不对称的。忽略这一点可能会导致数据移动成本的增加。因此,为了充分利用这些节点的潜力并降低数据访问成本,对机器拓扑(即计算节点拓扑和节点之间的互连网络)有一个完整的视图变得至关重要。此外,并行应用行为对如何有效地利用机器具有重要作用。在本文中,我们提出了一种分层负载平衡方法来提高应用程序在并行多核系统上的性能。我们介绍了一个拓扑感知负载均衡器NucoLB,它专注于重新分配工作,同时降低计算节点之间和节点内部的通信成本。NucoLB在其平衡决策中考虑了NUMA多核计算节点上存在的非对称内存访问成本、互连网络开销和应用程序通信模式。我们使用charm++并行运行时系统实现了NucoLB,并对其性能进行了评估。结果表明,与最先进的负载平衡器在三台不同的NUMA并行机器上相比,我们的负载平衡器将性能提高了20%。
Multi-core compute nodes with non-uniform memory access (NUMA) are now a common architecture in the assembly of large-scale parallel machines. On these machines, in addition to the network communication costs, the memory access costs within a compute node are also asymmetric. Ignoring this can lead to an increase in the data movement costs. Therefore, to fully exploit the potential of these nodes and reduce data access costs, it becomes crucial to have a complete view of the machine topology (i.e. the compute node topology and the interconnection network among the nodes). Furthermore, the parallel application behavior has an important role in determining how to utilize the machine efficiently. In this paper, we propose a hierarchical load balancing approach to improve the performance of applications on parallel multi-core systems. We introduce NucoLB, a topology-aware load balancer that focuses on redistributing work while reducing communication costs among and within compute nodes. NucoLB takes the asymmetric memory access costs present on NUMA multi-core compute nodes, the interconnection network overheads, and the application communication patterns into account in its balancing decisions. We have implemented NucoLB using the Charm++ parallel runtime system and evaluated its performance. Results show that our load balancer improves performance up to 20% when compared to state-of-the-art load balancers on three different NUMA parallel machines.