A Novel Shared-Memory Thread-Pool Implementation for Hybrid Parallel CFD Solvers

A Novel Shared-Memory Thread-Pool Implementation for Hybrid Parallel CFD Solvers
复制标题

混合并行 CFD 求解器的新型共享内存线程池实现

DOI:
10.1007/978-3-642-23397-5_18
复制
发表时间:
2011
期刊:
SIGMETRICS Perform. Evaluation Rev.
影响因子:
--
通讯作者:
C. Simmendinger
C. Simmendinger
中科院分区:
--
文献类型:
--
作者:
J. Jägersküpper;C. Simmendinger

文献摘要

被引文献

相似文献

非结构网格的计算流体动力学(CFD)求解器TAU在欧洲航空航天工业中得到了广泛的应用。Tau运行在具有数千个核心的高性能计算(HPC)集群上,使用基于MPI的域分解。为了更有效地利用现有的多核CPU,并为多核时代的TAU做好准备,在TAU的一个求解器中添加了共享内存并行化,以获得混合并行化:基于MPI的域分解+域的多线程处理。 对于所考虑的基于边的求解器,由于Amdahl陷阱,通过OpenMP for Directions的简单的基于循环的方法将不能提供所需的加速。开发了一种更复杂的、基于线程池的共享内存并行化,它允许通过自动和动态负载平衡进行轻松的线程同步。 在本文中,我们描述了这种共享内存并行化背后的概念,并解释了域的多线程计算是如何工作的。文中给出了它在TAU中的实现细节以及一些初步的性能结果。我们强调,这一概念不是TAU特有的。实际上,这种设计模式似乎非常通用,可以很好地应用于其他基于网格/网格/图形的代码。
The Computational Fluid Dynamics (CFD) solver TAU for unstructured grids is widely used in the European aerospace industry. TAU runs on High-Performance Computing (HPC) clusters with several thousands of cores using MPI-based domain decomposition. In order to make more efficient use of current multi-core CPUs and to prepare TAU for the many-core era, a shared-memory parallelization has been added to one of TAU's solver to obtain a hybrid parallelization: MPI-based domain decomposition plus multi-threaded processing of a domain. For the edge-based solver considered, a simple loop-based approach via OpenMP FOR directives would - due to the Amdahl trap - not deliver the required speed-up. A more sophisticated, thread-pool-based sharedmemory parallelization has been developed which allows for a relaxed thread synchronization with automatic and dynamic load balancing. In this paper we describe the concept behind this shared-memory parallelization, we explain how the multi-threaded computation of a domain works. Some details of its implementation in TAU as well as some first performance results are presented. We emphasize that the concept is not TAU-specific. Actually, this design pattern appears to be very generic and may well be applied to other grid/mesh/graph-based codes.