First Impressions of the NVIDIA Grace CPU Superchip and NVIDIA Grace Hopper Superchip for Scientific Workloads

First Impressions of the NVIDIA Grace CPU Superchip and NVIDIA Grace Hopper Superchip for Scientific Workloads
复制标题

针对科学工作负载的 NVIDIA Grace CPU 超级芯片和 NVIDIA Grace Hopper 超级芯片的第一印象

DOI:
10.1145/3636480.3637097
复制
发表时间:
2024
期刊:
ACM
影响因子:
--
通讯作者:
Harrison, Robert J.
Harrison, Robert J.
中科院分区:
--
文献类型:
--
作者:
Simakov, Nikolay A.;Jones, Matthew D.;Furlani, Thomas R.;Siegmann, Eva;Harrison, Robert J.

文献摘要

参考文献

相似文献

NVIDIA Grace CPU超级芯片和NVIDIA Grace Hopper超级芯片的工程样本使用不同的基准和科学应用进行了测试。基准包括HPCC和HPCG。真正的基于应用的基准测试包括AI-Benchmark-Alpha(TensorFlow基准测试)、Gromacs、OpenFOAM和ROMS。性能与多个英特尔、AMD、ARM CPU和几个采用NVIDIA GPU系统的x86进行了比较。根据TDP值进行了简要的能效评估。我们发现,在HPCC基准测试中,Grace的每核性能与AMD米兰内核相似或更快,而较高的内核数量往往使NVIDIA Grace CPU超级芯片具有与高带宽内存的Intel Sapphire Rapids相似的每节点性能:矩阵乘法(17%)和FFT(6%)更慢,Linpack(9%)更快)。在科学应用中,NVIDIA Grace CPU超级芯片的性能在Gromacs中慢6%到18%,在OpenFOAM中快7%,在只读存储器中介于Intel Sapphire Rapids的HBM和DDR模式之间。与任何经过测试的x86-NVIDIA GPU系统相比,Gromacs中的CPU-GPU组合性能显著更快(快20%到117%)。总体而言,新的NVIDIA Grace Hopper超级芯片和NVIDIA Grace CPU超级芯片是HPC中心的高性能和最有可能的节能解决方案。
The engineering samples of the NVIDIA Grace CPU Superchip and NVIDIA Grace Hopper Superchips were tested using different benchmarks and scientific applications. The benchmarks include HPCC and HPCG. The real application-based benchmark includes AI-Benchmark-Alpha (a TensorFlow benchmark), Gromacs, OpenFOAM, and ROMS. The performance was compared to multiple Intel, AMD, ARM CPUs and several x86 with NVIDIA GPU systems. A brief energy efficiency estimate was performed based on TDP values. We found that in HPCC benchmark tests, the per-core performance of Grace is similar to or faster than AMD Milan cores, and the high core count often allows NVIDIA Grace CPU Superchip to have per-node performance similar to Intel Sapphire Rapids with High Bandwidth Memory: slower in matrix multiplication (by 17%) and FFT (by 6%), faster in Linpack (by 9%)). In scientific applications, the NVIDIA Grace CPU Superchip performance is slower by 6% to 18% in Gromacs, faster by 7% in OpenFOAM, and right between HBM and DDR modes of Intel Sapphire Rapids in ROMS. The combined CPU-GPU performance in Gromacs is significantly faster (by 20% to 117% faster) than any tested x86-NVIDIA GPU system. Overall, the new NVIDIA Grace Hopper Superchip and NVIDIA Grace CPU Superchip Superchip are high-performance and most likely energy-efficient solutions for HPC centers.
DOI: 10.1002/jcc.24030
发表时间: 2015-10-05
影响因子: 3
作者:
Kutzner, Carsten;Pall, Szilard;Fechner, Martin;Esztermann, Ansgar;de Groot, Bert L.;Grubmueller, Helmut
通讯作者: Grubmueller, Helmut
DOI: 10.1002/cpe.5110
发表时间: 2019
期刊: Concurrency and Computation: Practice and Experience
影响因子: --
作者:
Simon McIntosh;J. Price;Tom Deakin;Andrei Poenaru
通讯作者: Andrei Poenaru
使用 XDMoD 促进 XSEDE 操作、规划和分析
DOI: 10.1145/2484762.2484763
发表时间: 2013
期刊: Proceedings of the Conference on Extreme Science and Engineering Discovery Environment: Gateway to Discovery (XSEDE '13
影响因子: --
作者:
Furlani, Thomas R.;Schneider, Barry L.;Jones, Matthew D.;Towns, John;Hart, David L.;Gallo, Steven M.;DeLeon, Robert L.;Lu, Charng-Da;Ghadersohi, Amin;Gentner, Ryan J.
通讯作者: Gentner, Ryan J.
CloudBank:简化计算机科学研究和教育云访问的托管服务
DOI: --
发表时间: 2021
期刊: Practice and Experience in Advanced Research Computing
影响因子: --
作者:
Michael L. Norman;Vince Kellen;Shava Smallen;Brian DeMeulle;Shawn M. Strande;Ed Lazowska;Naomi Alterman;R. Fatland;S. Stone;Amanda Tan;K. Yelick;Eric Van Dusen;James Mitchell
通讯作者: James Mitchell
DOI: 10.1145/3581576.3581618
发表时间: 2023-02
期刊: Proceedings of the HPC Asia 2023 Workshops
影响因子: --
作者:
N. Simakov;R. L. Deleon;Joseph P. White;Matthew D. Jones;T. Furlani;E. Siegmann;R. J. Harrison
通讯作者: N. Simakov;R. L. Deleon;Joseph P. White;Matthew D. Jones;T. Furlani;E. Siegmann;R. J. Harrison