High-performance lattice QCD for multi-core based parallel systems using a cache-friendly hybrid threaded-MPI approach

High-performance lattice QCD for multi-core based parallel systems using a cache-friendly hybrid threaded-MPI approach
复制标题

使用缓存友好的混合线程 MPI 方法,用于基于多核的并行系统的高性能晶格 QCD

DOI:
--
复制
发表时间:
2011
期刊:
2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子:
--
通讯作者:
P. Dubey
P. Dubey
中科院分区:
--
文献类型:
--
作者:
M. Smelyanskiy;K. Vaidyanathan;Jee Choi;B. Joó;J. Chhugani;M. Clark;P. Dubey

文献摘要

被引文献

相似文献

晶格量子色动力学 (LQCD) 是一个计算上具有挑战性的问题,需要在 SU(3) 规范场存在的情况下求解离散狄拉克方程。其关键运算是矩阵向量乘积,称为 Dslash 运算符。我们开发了一种新颖的多核架构友好型 Wilson-Dslash 运算符实现,它在英特尔® 至强® 处理器 X5680 上提供 75 Gflops(单精度),对于适合末级缓存的数据集实现 60% 的计算效率。对于大于最后一级缓存的数据集,此性能下降至 50 Gflops。在相同的硬件平台上运行时,我们的性能比 Chroma 软件套件的知名实现高 2-3 倍。本文报道的 LQCD 的新颖实现基于最近发布的 3.5D 空间和 4.5D 时间切片方案。两种阻塞方案都显着降低了 LQCD 外部存储器带宽要求,从而提供了更受计算限制的实现。随着计算触发器和外部存储器带宽之间的差距不断扩大,我们的方案的性能优势将变得更加显着。我们展示了我们的实现非常好的集群级可扩展性:对于 32 x 256 个站点的网格,当强大扩展到 128 个节点系统(总共 1536 个核心)时,我们实现了超过 4 Tflops。对于相同的晶格尺寸,完整的共轭梯度 Wilson-Dslash 算子可实现 2.95 Tflops。
Lattice Quantum Chromo-dynamics (LQCD) is a computationally challenging problem that solves the discretized Dirac equation in the presence of an SU(3) gauge field. Its key operation is a matrix-vector product, known as the Dslash operator. We have developed a novel multicore architecture-friendly implementation of the Wilson-Dslash operator which delivers 75 Gflops (single-precision) on an Intel® Xeon® Processor X5680 achieving 60% computational efficiency for datasets that fit in the last-level cache. For datasets larger than the last-level cache, this performance drops to 50 Gflops. Our performance is 2-3X higher than a well-known implementation from the Chroma software suite when running on the same hardware platform. The novel implementation of LQCD reported in this paper is based on recently published the 3.5D spatial and 4.5D temporal tiling schemes. Both blocking schemes significantly reduce LQCD external memory bandwidth requirements, delivering a more compute-bound implementation. The performance advantage of our schemes will become more significant as the gap between compute flops and external memory bandwidth continues to grow. We demonstrate very good cluster-level scalability of our implementation: for a lattice of 32 x 256 sites, we achieve over 4 Tflops when strong-scaled to a 128 node system (1536 cores total). For the same lattice size, a full Conjugate Gradients Wilson-Dslash operator, achieves 2.95 Tflops.