Iterative methods with mixed-precision preconditioning for ill-conditioned linear systems in multiphase CFD simulations

Iterative methods with mixed-precision preconditioning for ill-conditioned linear systems in multiphase CFD simulations
复制标题

多相 CFD 仿真中病态线性系统的混合精度预处理迭代方法

DOI:
10.1109/scala54577.2021.00006
复制
发表时间:
2021
期刊:
Proceedings of 12th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems (ScalA)
影响因子:
--
通讯作者:
Onodera N.
Onodera N.
中科院分区:
--
文献类型:
--
作者:
Ina T.;Idomura Y.;Imamura T.;Yamashita S.;Onodera N.

文献摘要

相似文献

针对多相热工水力CFD程序JUPITER [1]中的预处理共轭梯度(P-CG)求解器和多重网格预处理共轭梯度(MGCG)求解器,开发了一种新的基于迭代精化(IR)方法的混合精度预处理器。在IR预处理器中,使用块Jacobi方法近似求解线性系统,该方法使用具有2D瓦片的精细块分解进行优化,以促进SIMD操作和软件流水线。混合精度计算的目的是避免在计算系数相差很大的病态矩阵时舍入误差的影响。线性系统被归一化,以便可以在FP 16的动态范围内计算。所有数据都存储在FP 16中以减少内存访问,而所有计算都在FP 32中执行。混合算法FP 16/32保持了与FP 32相似的收敛性能,而计算性能接近FP 16。FP 16/32实现的鲁棒性通过扫描具有不同有效位长度的数据格式来证明。所开发的求解器在Fugaku(A64 FX)上进行了优化,并应用于JUPITER中具有900亿自由度的病态矩阵。具有新的IR预处理器的P-CG和MGCG求解器显示出高达8,000个CPU的出色的强大扩展能力,其中在Oakforest-PACS(KNL)上的传统求解器分别实现了5.7倍和3.4倍的加速比。
A new mixed-precision preconditioner based on the iterative refinement (IR) method is developed for the preconditioned conjugate gradient (P-CG) solver and the multigrid preconditioned conjugate gradient (MGCG) solver in the multiphase thermal-hydraulic CFD code JUPITER [1]. In the IR preconditioner, linear systems are approximately solved using a block Jacobi method, which is optimized using fine block decomposition with 2D tiles to facilitate SIMD operations and software pipelining. Mixed-precision computing is designed to avoid influences of the roundoff errors in computing ill-conditioned matrices with extreme contrast of coefficients. Linear systems are normalized so that it can be computed within the dynamic range of FP16. All data is stored in FP16 to reduce memory access, while all computation is performed in FP32. The hybrid FP16/32 implementation keeps the similar convergence property as FP32, while the computational performance is close to FP16. The robustness of the FP16/32 implementation was demonstrated by scanning data formats with different bit lengths of the significand. The developed solvers are optimized on Fugaku (A64FX), and applied to ill-conditioned matrices with 90 billion DOFs in JUPITER. The P-CG and MGCG solvers with the new IR preconditioner show excellent strong scaling up to 8,000 CPUs, where 5.7 × and 3.4 × speedups are respectively achieved from the conventional solvers on Oakforest-PACS (KNL).