A new metric for ranking high-performance computing systems
A new metric for ranking high-performance computing systems
复制标题
用于对高性能计算系统进行排名的新指标
DOI:
10.1093/nsr/nwv084
复制
发表时间:
2016
影响因子:
20.6
通讯作者:
P. Luszczek
中科院分区:
文献类型:
--
作者:
J. Dongarra;M. Heroux;P. Luszczek
Performance benchmarks play an important role in various stages of hardware development and use. During development, hardware designers use benchmarks as proxies for full-fledged applications because prototype systems have a limited set of software tools that the applications require for compilation and execution. At procurement time, benchmarks serve as an assurance test that establishes the new system’s viability and fulfillment of contractual obligations between the system integrator and the customer. And finally, benchmarking during the system’s daily use can ensure proper operation of a computer installation, exposing any potential problems to the system’s administrators, and gives the users an estimate of what their applications of interest can achieve without requiring the effort of building those applications, their software dependences, and loading the necessary input data that may initially reside offsite. The LINPACK benchmark [1] has been in continuous use since the 1980s. It was born out of necessity in the 1970s, when it was used to quickly test the performance of vector subroutines—which could serve as a good approximation of performance for the rest of the LINPACK library. Because of the nature of the implementation, the benchmark could also be used as a first-order approximation of other codes, partially due to the well-balanced hardware that offered plentiful bandwidth for every floating-point operation. Over the years, Moore’s law eroded the compute-tobandwidth balance, resulting in a memory wall. Today, this wall can be exhibited by the Intel Haswell (Xeon E7 v3 8900) processors, which feature 18 cores—each equipped with dual floatingpoint units (FPUs) with AVX vectors capable of fused multiply-add (FMA) instructions and clocked at 2.5 GHz (before TurboBoost)—bringing to bear, altogether, over 700 Gflop/s worth of peak performance—over 90% of which may be realized in a well-tuned matrix–matrix multiplication from a vendor library (MKL). At the same time, the memory controller for that processor has a theoretical maximum of 25 GB/s. The exact number of floating operations per every byte of bandwidth is 28.8—a far different ratio than the 1 flop per byte, which was the design point in the 1990s when the TOP500 list started ranking the supercomputers using the scalable version of the LINPACK benchmark. To reassess the application needs in this new and drastically different hardware regime, it is worthwhile to look at the computational simulations that drive national interest. Many of these simulations involve heat diffusion, electromagnetics, and fluid dynamics. Unlike LINPACK, which tests raw floatingpoint performance and its delivery through the BLAS API, these real-world applications rely on partial differential equations (PDEs) that govern the continuous representations of the physical quantities such as particle speed, momentum, etc. These PDEs involve sparse (not dense) matrices that represent the 3Dl embedding of the discretization mesh. While the size of the sparse data fills the available memory to accommodate the simulation models of interest, most of the optimization techniques that help achieve close to peak performance in dense matrix calculations are only marginally useful in the context of sparse matrices originating from PDEs. Our new benchmark, called highperformance conjugate gradients (HPCG)(further information is available at www. hpcg-benchmark. org), is based on Mantevo collection’s HPCCG code base [2], but aims to go beyond its originator and represent the calculations that commonly occur during the numerical solution of PDEs in modern stateof-the-art solvers. To that end, HPCG is …