Multi-block/multi-core SSOR preconditioner for the QCD quark solver for K computer

Multi-block/multi-core SSOR preconditioner for the QCD quark solver for K computer
复制标题

K计算机QCD夸克求解器的多块/多核SSOR预处理器

DOI:
10.22323/1.164.0188
复制
发表时间:
2012
期刊:
arXiv: High Energy Physics - Lattice
影响因子:
--
通讯作者:
T. Yoshie
T. Yoshie
中科院分区:
--
文献类型:
--
作者:
T. Boku;K. Ishikawa;Y. Kuramashi;K. Minami;Y. Nakamura;F. Shoji;D. Takahashi;M. Terai;A. Ukawa;T. Yoshie

文献摘要

被引文献

相似文献

本文研究了格点QCD三叶草-费米子解算器在K计算机上的算法优化和性能调整。我们实现了L“uscher的SAP预条件子分块,其中在一个节点中的格块进一步划分为几个子块,以提取足够的并行度为8核CPU的SPARC 64 $^{\mathrm{TM}}$ VIIIfx的K计算机。为了实现更好的收敛性能,我们使用对称连续超松弛(SSOR)迭代与{\it局部字典序}排序的子块在获得块逆。SAP预处理器包含在嵌套BiCGStab求解器的单精度BiCGStab求解器中。计算内核的单精度部分仅用面向SIMD的内部函数编写,以实现K计算机上的\n的最佳性能。我们在三种晶格尺寸上对单精度BiCGStab求解器进行基准测试:$12^3\times 24$,$24^3\times 48$和$48^3\times 96$,将节点中的局部晶格尺寸固定为$6^3\times 12$。我们观察到一个理想的弱伸缩性能从16个节点到4096个节点。计算内核的性能超过50%的效率,单精度BiCGstab具有$\sim26%的持续效率。
We study the algorithmic optimization and performance tuning of the Lattice QCD clover-fermion solver for the K computer. We implement the L\"uscher's SAP preconditioner with sub-blocking in which the lattice block in a node is further divided to several sub-blocks to extract enough parallelism for the 8-core CPU SPARC64$^{\mathrm{TM}}$ VIIIfx of the K computer. To achieve a better convergence property we use the symmetric successive over-relaxation (SSOR) iteration with {\it locally-lexicographical} ordering for the sub-blocks in obtaining the block inverse. The SAP preconditioner is included in the single precision BiCGStab solver of the nested BiCGStab solver. The single precision part of the computational kernel are solely written with the SIMD oriented intrinsics to achieve the best performance of the \SPARC on the K computer. We benchmark the single precision BiCGStab solver on the three lattice sizes: $12^3\times 24$, $24^3\times 48$ and $48^3\times 96$, with fixing the local lattice size in a node at $6^3\times 12$. We observe an ideal weak-scaling performance from 16 nodes to 4096 nodes. The performance of a computational kernel exceeds 50% efficiency, and the single precision BiCGstab has $\sim26% susutained efficiency.