A new metric for ranking high-performance computing systems

A new metric for ranking high-performance computing systems
复制标题

用于对高性能计算系统进行排名的新指标

DOI:
10.1093/nsr/nwv084
复制
发表时间:
2016
影响因子:
20.6
通讯作者:
P. Luszczek
P. Luszczek
中科院分区:
综合性期刊1区
文献类型:
--
作者:
J. Dongarra;M. Heroux;P. Luszczek

文献摘要

被引文献

相似文献

性能基准测试在硬件开发和使用的各个阶段都扮演着重要的角色。在开发过程中,硬件设计人员使用基准测试作为成熟应用程序的代理,因为原型系统具有应用程序编译和执行所需的有限软件工具集。在采购时,基准作为保证测试,确定新系统的可行性和系统集成商和客户之间的合同义务的履行。最后,在系统的日常使用期间进行基准测试可以确保计算机安装的正确操作,向系统管理员暴露任何潜在的问题,并向用户提供他们感兴趣的应用程序可以实现的估计,而不需要构建这些应用程序、它们的软件依赖关系和加载最初可能驻留在场外的必要输入数据。自20世纪80年代以来,LINPACK基准[1]一直在持续使用。它诞生于20世纪70年代,当时它被用来快速测试向量子例程的性能——这可以作为LINPACK库其余部分性能的一个很好的近似。由于实现的性质,基准测试还可以用作其他代码的一阶近似值,部分原因是平衡良好的硬件为每个浮点操作提供了充足的带宽。多年来,摩尔定律侵蚀了计算机与带宽的平衡,导致了内存墙。今天,这道墙可以通过Intel Haswell (Xeon E7 v3 8900)处理器展示出来,它有18个内核,每个内核都配备了双浮点单元(fpu), AVX矢量能够融合乘加(FMA)指令,时钟频率为2.5 GHz(在TurboBoost之前),总共可以实现超过700 Gflop/s的峰值性能,其中90%可以通过供应商库(MKL)的优化矩阵-矩阵乘法实现。同时,该处理器的内存控制器的理论最大值为25gb /s。每字节带宽的浮点运算的确切数量是28.8,这与每字节1个flop的比率大不相同,这是20世纪90年代TOP500列表开始使用LINPACK基准的可伸缩版本对超级计算机进行排名时的设计点。为了重新评估这种全新的、完全不同的硬件系统中的应用程序需求,有必要研究一下驱动国家利益的计算模拟。其中许多模拟涉及热扩散、电磁学和流体动力学。与LINPACK不同的是,LINPACK通过BLAS API测试原始浮点性能及其交付,这些现实世界的应用程序依赖于控制物理量(如粒子速度、动量等)的连续表示的偏微分方程(pde)。这些偏微分方程涉及稀疏(不密集)矩阵,表示离散网格的3Dl嵌入。虽然稀疏数据的大小填充了可用内存,以容纳感兴趣的模拟模型,但大多数有助于在密集矩阵计算中实现接近峰值性能的优化技术在源自pde的稀疏矩阵上下文中只是略微有用。我们的新基准,称为高性能共轭梯度(HPCG)(更多信息可在www。hpcg-benchmark。org),基于Mantevo collection的HPCCG代码库[2],但旨在超越其创始人,并代表在现代最先进的解算器中通常发生的偏微分方程数值解的计算。为此,HPCG…
Performance benchmarks play an important role in various stages of hardware development and use. During development, hardware designers use benchmarks as proxies for full-fledged applications because prototype systems have a limited set of software tools that the applications require for compilation and execution. At procurement time, benchmarks serve as an assurance test that establishes the new system’s viability and fulfillment of contractual obligations between the system integrator and the customer. And finally, benchmarking during the system’s daily use can ensure proper operation of a computer installation, exposing any potential problems to the system’s administrators, and gives the users an estimate of what their applications of interest can achieve without requiring the effort of building those applications, their software dependences, and loading the necessary input data that may initially reside offsite. The LINPACK benchmark [1] has been in continuous use since the 1980s. It was born out of necessity in the 1970s, when it was used to quickly test the performance of vector subroutines—which could serve as a good approximation of performance for the rest of the LINPACK library. Because of the nature of the implementation, the benchmark could also be used as a first-order approximation of other codes, partially due to the well-balanced hardware that offered plentiful bandwidth for every floating-point operation. Over the years, Moore’s law eroded the compute-tobandwidth balance, resulting in a memory wall. Today, this wall can be exhibited by the Intel Haswell (Xeon E7 v3 8900) processors, which feature 18 cores—each equipped with dual floatingpoint units (FPUs) with AVX vectors capable of fused multiply-add (FMA) instructions and clocked at 2.5 GHz (before TurboBoost)—bringing to bear, altogether, over 700 Gflop/s worth of peak performance—over 90% of which may be realized in a well-tuned matrix–matrix multiplication from a vendor library (MKL). At the same time, the memory controller for that processor has a theoretical maximum of 25 GB/s. The exact number of floating operations per every byte of bandwidth is 28.8—a far different ratio than the 1 flop per byte, which was the design point in the 1990s when the TOP500 list started ranking the supercomputers using the scalable version of the LINPACK benchmark. To reassess the application needs in this new and drastically different hardware regime, it is worthwhile to look at the computational simulations that drive national interest. Many of these simulations involve heat diffusion, electromagnetics, and fluid dynamics. Unlike LINPACK, which tests raw floatingpoint performance and its delivery through the BLAS API, these real-world applications rely on partial differential equations (PDEs) that govern the continuous representations of the physical quantities such as particle speed, momentum, etc. These PDEs involve sparse (not dense) matrices that represent the 3Dl embedding of the discretization mesh. While the size of the sparse data fills the available memory to accommodate the simulation models of interest, most of the optimization techniques that help achieve close to peak performance in dense matrix calculations are only marginally useful in the context of sparse matrices originating from PDEs. Our new benchmark, called highperformance conjugate gradients (HPCG)(further information is available at www. hpcg-benchmark. org), is based on Mantevo collection’s HPCCG code base [2], but aims to go beyond its originator and represent the calculations that commonly occur during the numerical solution of PDEs in modern stateof-the-art solvers. To that end, HPCG is …