Lattice-CSC: Optimizing and Building an Efficient Supercomputer for Lattice-QCD and to Achieve First Place in Green500

Lattice-CSC: Optimizing and Building an Efficient Supercomputer for Lattice-QCD and to Achieve First Place in Green500
复制标题

Lattice-CSC:为Lattice-QCD优化和构建高效超级计算机并在Green500中获得第一名

DOI:
10.1007/978-3-319-20119-1_14
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
O. Philipsen
O. Philipsen
中科院分区:
--
文献类型:
--
作者:
D. Rohr;M. Bach;G. Nešković;V. Lindenstruth;C. Pinke;O. Philipsen

文献摘要

参考文献

被引文献

相似文献

在过去的几十年里,超级计算机已经成为科学和工业的必需品。巨大的数据中心消耗大量的电力,我们正处于更新,更快的计算机必须不再比它们的前辈消耗更多的电力的时刻。用户对计算能力的需求没有以任何方式下降,这一事实导致了对exaflop系统可行性的研究。具有高效加速器(如GPU)的异构集群是提高效率的一种方法。我们提出了新的L-CSC集群,商品硬件计算集群专用于格QCD模拟在GSI研究设施。L-CSC采用多GPU设计,每个节点有四个FirePro S9150 GPU,提供320 GB/s的内存带宽和2.6 TFLOPS的峰值性能。高带宽使其非常适合内存受限的LQCD计算,而多GPU设计确保了上级能效。2014年11月Green 500名单授予L-CSC世界上最节能的超级计算机,在Linpack基准测试中达到5270 MFLOPS/W。本文介绍了对我们的Linpack实现HPL-GPU的优化以及帮助L-CSC达到这一基准的其他能效改进。它描述了我们的方法,一个准确的Green 500功率测量,并揭示了目前的测量方法的一些问题。最后,对格点QCD在L-CSC中的应用进行了综述。
In the last decades, supercomputers have become a necessity in science and industry. Huge data centers consume enormous amounts of electricity and we are at a point where newer, faster computers must no longer drain more power than their predecessors. The fact that user demand for compute capabilities has not declined in any way has led to studies of the feasibility of exaflop systems. Heterogeneous clusters with highly-efficient accelerators such as GPUs are one approach to higher efficiency. We present the new L-CSC cluster, a commodity hardware compute cluster dedicated to Lattice QCD simulations at the GSI research facility. L-CSC features a multi-GPU design with four FirePro S9150 GPUs per node providing 320 GB/s memory bandwidth and 2.6 TFLOPS peak performance each. The high bandwidth makes it ideally suited for memory-bound LQCD computations while the multi-GPU design ensures superior power efficiency. The November 2014 Green500 list awarded L-CSC the most power-efficient supercomputer in the world with 5270 MFLOPS/W in the Linpack benchmark. This paper presents optimizations to our Linpack implementation HPL-GPU and other power efficiency improvements which helped L-CSC reach this benchmark. It describes our approach for an accurate Green500 power measurement and unveils some problems with the current measurement methodology. Finally, it gives an overview of the Lattice QCD application on L-CSC.
QCDOC:用于紧耦合计算的 10 Teraflops 计算机
DOI: --
发表时间: 2004
期刊: Proceedings of the ACM/IEEE SC2004 Conference
影响因子: --
作者:
P. Boyle;Dong Chen;N. Christ;M. Clark;Saul D. Cohen;Z. Dong;A. Gara;B. Joó;C. Jung;L. Levkova;X. Liao;G. Liu;R. Mawhinney;S. Ohta;K. Petrov;T. Wettig;A. Yamaguchi;C. Cristian
通讯作者: C. Cristian
N f = 2 QCD 中威尔逊费米子的 Roberge-Weiss 跃迁的性质
DOI: --
发表时间: 2014
期刊:
影响因子: --
作者:
F. Cuteri;C. Czaban;O. Philipsen;C. Pinke;A. Sciarra
通讯作者: A. Sciarra
DOI: --
发表时间: 2013
期刊: 2013 21st Euromicro International Conference on Parallel, Distributed, and Network-Based Processing
影响因子: --
作者:
M. Bach;J. Cuveland;H. Ebermann;D. Eschweiler;J. Gerhard;S. Kalcher;M. Kretz;V. Lindenstruth;H. Ludde;Manfred Pollok;D. Rohr
通讯作者: D. Rohr
使用缓存友好的混合线程 MPI 方法,用于基于多核的并行系统的高性能晶格 QCD
DOI: --
发表时间: 2011
期刊: 2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子: --
作者:
M. Smelyanskiy;K. Vaidyanathan;Jee Choi;B. Joó;J. Chhugani;M. Clark;P. Dubey
通讯作者: P. Dubey
GPU 上的格子 QCD 计算框架
DOI: 10.1109/ipdps.2014.112
发表时间: 2014
期刊: 2014 IEEE 28th International Parallel and Distributed Processing Symposium
影响因子: --
作者:
F. Winter;M. Clark;R. Edwards;B. Joó
通讯作者: B. Joó