Performance Tuning of a Lattice QCD Code on aN ode of the Kc omputer

Performance Tuning of a Lattice QCD Code on aN ode of the Kc omputer
复制标题

Kc 计算机节点上格子 QCD 码的性能调优

DOI:
--
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
M. Yokokawa
M. Yokokawa
中科院分区:
--
文献类型:
--
作者:
M. Terai;K. Ishikawa;Yoshinori Sugisaki;K. Minami;F. Shoji;Y. Nakamura;Y. Kuramashi;M. Yokokawa

文献摘要

参考文献

被引文献

相似文献

格子量子色动力学是基于强相互作用解决夸克和胶子之间动力学的第一原理计算。该计算是在离散化为格子的四维时空上进行的,需要对Wilson-Dirac方程导出的稀疏矩阵进行大量的求逆。在本研究中,格子 QCD 代码、LDDHMC 使用域分解 HMC 算法和混合精度 BiCGStab 求解器来求解线性方程。该方案是嵌套的,由内求解器和外求解器组成。外部求解器是双精度 BiCGStab 的计算。内部求解器是单精度 BiCGStab 的预处理计算,并由 Luscher 的 SAP 进行预处理。此外,利用SSOR改进了SAP小块的计算。为了提高性能,我们从应用程序代码中的SSOR例程中提取了三个内核代码,并通过分析器分析了内核的瓶颈。基于分析,我们得到了以下几点问题:a) SIMD 指令速率,b) 整数 L1D 缓存未命中,c) 浮点 L1D 缓存未命中,d) 指令调度,e) 屏障同步。结果,调整将内核的峰值性能从 kernel-1 中的 23.2% 提高到 38.1%,将 kernel-2 中的峰值性能从 24.3% 提高到 38.0%,将 kernel-3 中的峰值性能从 23.6% 提高到 44.9%。芯片的峰值性能在 kernel-1 中为 29.5%,在 kernel-2 中为 30.9%,在 kernel-3 中为 37.8%。结果显示了通过分析和调整来改进编译器的有效性。
Lattice QCD is first principle calculation to solve the dynamics between quarks and gluons based on strong interaction. The calculation is performed on four dimensional space-time which is discretized to lattice, and requires a huge amount of inversion of the sparse matrix derived from Wilson-Dirac equation. In this study, Lattice QCD code, LDDHMC uses domain decomposition HMC algorithm with mixed precision BiCGStab solver for the linear equation. This scheme is nested, consists of inner solver and outer solver. The outer solver is calculation of BiCGStab with double precision. The inner solver is preconditioning calculation of BiCGStab with single precision and is preconditioned by the Luscher's SAP. Furthermore, the calculation for the small block of SAP is improved with SSOR. To improve the performance we extracted three kernel codes from the SSOR routine in the application codes, and analyzed bottlenecks for the kernels by profiler. Based on the profiling we obtained the problems for following points: a) SIMD instruction rate, b) integer L1D cache misses, c) floating-point L1D cache misses, d) instruction scheduling, e) barrier synchronization. As a result, the tuning improves the peak performance a core from 23.2% to 38.1% in the kernel-1, from 24.3% to 38.0% in the kernel-2, from 23.6% to 44.9% in the kernel-3. The peak performance a chip is 29.5% in the kernel-1, 30.9% in the kernel-2, 37.8% in the kernel-3. The results show effectiveness for improvement of the compiler by profiling and tuning.
晶格 QCD 模拟作为 HPC 挑战
DOI: --
发表时间: 2008
期刊: High-Performance Computing, Lecture Notes in Computer Science 4759
影响因子: --
作者:
J.;Wei;Jaume Garriga and Takahiro Tanaka;H. Okamoto;Nathalie Deruelle;Masaru Shibata;H. Higaki;Atsushi Nakamura
通讯作者: Atsushi Nakamura