Testing performance with and without Block Low Rank Compression in MUMPS and the new PaStiX 6.0 for JOREK nonlinear MHD simulations

Testing performance with and without Block Low Rank Compression in MUMPS and the new PaStiX 6.0 for JOREK nonlinear MHD simulations
复制标题

在 MUMPS 和用于 JOREK 非线性 MHD 模拟的新 PaStiX 6.0 中测试有或没有块低阶压缩的性能

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
M. Hoelzl
M. Hoelzl
中科院分区:
--
文献类型:
--
作者:
R. Nies;M. Hoelzl

文献摘要

被引文献

相似文献

在JOREK MHD代码中更新了MUMPS求解器的接口,以支持块低秩(BLR)压缩,并实现了新PaStiX求解器版本6的接口,以支持BLR。第一个测试是用JOREK进行的,它在每个时间步迭代求解一个大型稀疏矩阵系统。对于预处理,在代码中将直接求解器应用于子矩阵,此时应用BLR,结果总结在本报告中。对于一个简单的情况下,线性增长模式,结果与两个求解器看起来很有前途的内存消耗的百分之几十,得到了相当大的减少。在特定配置中已经看到了性能的直接提高。 在这个简单的测试和更真实的模拟中,BLR精度参数的选择被证明是至关重要的,由于时间有限,这些模拟仅使用MUMPS进行。更实际的测试显示,使用BLR时运行时间会增加,而使用更大的$blog $值时会减少。然而,当$太大时,GMRES迭代求解器不再达到收敛,因为在这种情况下预处理器变得太不准确。因此,关键是要使用尽可能大的$$,同时仍然达到收敛。今后还需要进行更多关于这一最佳值的试验。BLR还可以在特定情况下导致间接加速,此时由于内存消耗减少,可以在更少数量的计算节点上运行模拟。
The interface to the MUMPS solver was updated in the JOREK MHD code to support Block Low Rank (BLR) compression and an interface to the new PaStiX solver version 6 has been implemented supporting BLR as well. First tests were carried out with JOREK, which solves a large sparse matrix system iteratively in each time step. For the preconditioning, a direct solver is applied in the code to sub-matrices, and at this point BLR was applied with the results being summarized in this report. For a simple case with a linearly growing mode, results with both solvers look promising with a considerable reduction of the memory consumption by several ten percent was obtained. A direct increase in performance was seen in particular configurations already. The choice of the BLR accuracy parameter $epsilon$ proves to be critical in this simple test and also in more realistic simulations, which were carried out only with MUMPS due to the limited time available. The more realistic test showed an increase in run time when using BLR, which was mitigated when using larger values of $epsilon$. However, the GMRes iterative solver does not reach convergence anymore when $epsilon$ is too large, since the preconditioner becomes too inaccurate in that case. It is thus critical to use an $epsilon$ as large as possible, while still reaching convergence. More tests regarding this optimum will be necessary in the future. BLR can also lead to an indirect speed-up in particular cases, when the simulation can be run on a smaller number of compute nodes due to the reduced memory consumption.