Dynamic load balancing with enhanced shared-memory parallelism for particle-in-cell codes

Dynamic load balancing with enhanced shared-memory parallelism for particle-in-cell codes
复制标题

DOI:
10.1016/j.cpc.2020.107633
复制
发表时间:
2020-03
期刊:
Comput. Phys. Commun.
影响因子:
--
通讯作者:
Kyle G. Miller;Roman Lee;A. Tableman;A. Helm;R. Fonseca;V. Decyk;W. Mori
Kyle G. Miller;Roman Lee;A. Tableman;A. Helm;R. Fonseca;V. Decyk;W. Mori
中科院分区:
其他
文献类型:
--
作者:
Kyle G. Miller;Roman Lee;A. Tableman;A. Helm;R. Fonseca;V. Decyk;W. Mori

文献摘要

相似文献

为了进一步理解当今等离子体物理学中许多有趣的问题-包括基于等离子体的加速和由于量子电动力学效应而产生的磁重联-需要使用粒子单元(PIC)代码进行大规模动力学模拟。然而,这些模拟是非常苛刻的,要求当代PIC代码的设计,以有效地使用一个新的舰队的exascale计算架构。为此,必须解决跨计算节点的并行负载平衡的关键问题。我们讨论了动态负载平衡的实现,通过将模拟空间划分为许多小的、自包含的区域或“瓦片”,沿着共享内存(例如,OpenMP)并行性。负载平衡算法可以用于三种不同的拓扑结构,包括两个空间填充曲线。我们在代码Osiris中测试了此实现,并在均匀负载和严重负载不平衡的模拟上显示了较低的开销和改进的可扩展性。与其他负载平衡技术相比,我们的算法提供了一个数量级的改进,在模拟与严重的负载不平衡问题的并行可扩展性。
Furthering our understanding of many of today’s interesting problems in plasma physics – including plasma based acceleration and magnetic reconnection with pair production due to quantum electrodynamic effects – requires large-scale kinetic simulations using particle-in-cell (PIC) codes. However, these simulations are extremely demanding, requiring that contemporary PIC codes be designed to efficiently use a new fleet of exascale computing architectures. To this end, the key issue of parallel load balance across computational nodes must be addressed. We discuss the implementation of dynamic load balancing by dividing the simulation space into many small, self-contained regions or “tiles,” along with shared-memory (e.g., OpenMP) parallelism both over many tiles and within single tiles. The load balancing algorithm can be used with three different topologies, including two space-filling curves. We tested this implementation in the codeOsirisand show low overhead and improved scalability with OpenMP thread number on simulations with both uniform load and severe load imbalance. Compared to other load-balancing techniques, our algorithm gives order-of-magnitude improvement in parallel scalability for simulations with severe load imbalance issues.