Scalable Post-Processing of Large-Scale Numerical Simulations of Turbulent Fluid Flows

Scalable Post-Processing of Large-Scale Numerical Simulations of Turbulent Fluid Flows
复制标题

DOI:
10.3390/sym14040823
复制
发表时间:
2022-04
期刊:
Symmetry
影响因子:
--
通讯作者:
Christian Lagares;Wilson Rivera;G. Araya
Christian Lagares;Wilson Rivera;G. Araya
中科院分区:
其他
文献类型:
--
作者:
Christian Lagares;Wilson Rivera;G. Araya

文献摘要

被引文献

相似文献

军事、太空和高速民用应用将继续促进对可压缩、高速湍流边界层的新兴趣。更复杂的是,这些流呈现出复杂的计算挑战,从预处理到大规模数值模拟的执行和随后的后处理。在更高雷诺数下探索更复杂的几何形状将需要可扩展的后处理。现代给应用程序开发人员和科学家带来了越来越多样化和异构的计算硬件,这大大复杂化了性能可移植应用程序的开发。为了应对这些挑战,我们提出了Aquila,这是一个分布式、核外、性能可移植的大规模模拟后处理库。它旨在减轻领域专家编写针对具有强伸缩性能的异构高性能计算机的应用程序的负担。我们提供了两种实现,c++和Python;并展示了它们强大的扩展性能和在内核外操作时达到峰值内存带宽的60%和峰值文件系统带宽的98%的能力。我们还提出了通过利用傅里叶空间中的对称性来优化两点相关性的方法。所提出的设计的一个关键区别是包含了一个核外数据预取器,以提供文件在内存中可用性的错觉,从而在程序运行时中产生高达46%的改进。此外,我们还演示了高线程工作负载的并行效率超过70%。
Military, space, and high-speed civilian applications will continue contributing to the renewed interest in compressible, high-speed turbulent boundary layers. To further complicate matters, these flows present complex computational challenges ranging from the pre-processing to the execution and subsequent post-processing of large-scale numerical simulations. Exploring more complex geometries at higher Reynolds numbers will demand scalable post-processing. Modern times have brought application developers and scientists the advent of increasingly more diversified and heterogeneous computing hardware, which significantly complicates the development of performance-portable applications. To address these challenges, we propose Aquila, a distributed, out-of-core, performance-portable post-processing library for large-scale simulations. It is designed to alleviate the burden of domain experts writing applications targeted at heterogeneous, high-performance computers with strong scaling performance. We provide two implementations, in C++ and Python; and demonstrate their strong scaling performance and ability to reach 60% of peak memory bandwidth and 98% of the peak filesystem bandwidth while operating out of core. We also present our approach to optimizing two-point correlations by exploiting symmetry in the Fourier space. A key distinction in the proposed design is the inclusion of an out-of-core data pre-fetcher to give the illusion of in-memory availability of files yielding up to 46% improvement in program runtime. Furthermore, we demonstrate a parallel efficiency greater than 70% for highly threaded workloads.