Optimized Lattice QCD kernels for a Pentium 4 Cluster

Optimized Lattice QCD kernels for a Pentium 4 Cluster
复制标题

针对 Pentium 4 集群优化的 Lattice QCD 内核

DOI:
10.2172/954827
复制
发表时间:
2001
影响因子:
5.4
通讯作者:
Christopher L. McClendon
Christopher L. McClendon
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Christopher L. McClendon

文献摘要

被引文献

相似文献

不久,一个新的并行Pentium 4机器集群将在JLAB建立,以运行Lattice QCD计算。我讨论了优化的Lattice QCD例程的基本原理,以及奔腾4的特性如何使新的优化例程比普通C例程运行得更快。我描述了SU(3)线性代数例程中使用的优化策略,以及威尔逊-狄拉克算子的单节点和并行实现。最后,我显示单节点性能的并行版本的威尔逊-狄拉克算子的时间。
Soon, a new cluster of parallel Pentium 4 machines will be set up at JLAB to run Lattice QCD calculations. I discuss the rationale for optimized Lattice QCD routines, and how the features of the Pentium 4 enable new optimized routines to run much faster than normal C routines. I describe the optimization strategies used in SU(3) linear algebra routines, and in both single-node and parallel implementations of the Wilson-Dirac Operator. Finally, I show single node performance timings for the parallel version of the Wilson-Dirac operator.