Optimized Lattice QCD kernels for a Pentium 4 Cluster
Optimized Lattice QCD kernels for a Pentium 4 Cluster
复制标题
针对 Pentium 4 集群优化的 Lattice QCD 内核
DOI:
10.2172/954827
复制
发表时间:
2001
影响因子:
5.4
通讯作者:
Christopher L. McClendon
中科院分区:
文献类型:
--
作者:
Christopher L. McClendon
Soon, a new cluster of parallel Pentium 4 machines will be set up at JLAB to run Lattice QCD calculations. I discuss the rationale for optimized Lattice QCD routines, and how the features of the Pentium 4 enable new optimized routines to run much faster than normal C routines. I describe the optimization strategies used in SU(3) linear algebra routines, and in both single-node and parallel implementations of the Wilson-Dirac Operator. Finally, I show single node performance timings for the parallel version of the Wilson-Dirac operator.