Evaluating the Intel Skylake Xeon Processor for HPC Workloads

Evaluating the Intel Skylake Xeon Processor for HPC Workloads
复制标题

评估适用于 HPC 工作负载的 Intel Skylake Xeon 处理器

DOI:
10.1109/hpcs.2018.00064
复制
发表时间:
2018
期刊:
2018 International Conference on High Performance Computing & Simulation (HPCS)
影响因子:
--
通讯作者:
C. Hughes
C. Hughes
中科院分区:
--
文献类型:
--
作者:
S. Hammond;C. Vaughan;C. Hughes

文献摘要

被引文献

相似文献

尽管在将科学应用程序移植到新架构(如计算优化的图形处理器、众核处理器/加速器,甚至专用功能单元)方面取得了重大进展,但绝大多数科学计算仍在高性能商用服务器处理器上执行。即使在已经移植到新架构的应用程序中,频繁的串行部分仍然需要强大的服务器级处理器内核来尽可能快地计算。在本文中,我们报告了一组评估英特尔最新Skylake Xeon服务器处理器的基准测试研究。Skylake代表了Xeon产品线的重大变化,具有更宽的SIMD向量单元,重新设计的缓存架构以及更多的内存通道。更宽的矢量单元为一些计算密集型应用程序提供了2倍的改进,并且组合的内存更改可以提供接近2倍的内存带宽。我们在几个与HPC相关的小型应用程序和基准测试中评估了这些新的硬件功能,包括STREAM、LULESH、XSBench、HPCG和SW 4Lite。与上一代Haswell处理器内核相比,新的硬件功能在HPC基准代码上提供了高达1.8倍的加速,为依赖此类计算节点的更广泛的HPC应用提供了更大的实用性。
Despite significant advances in the porting of scientific applications to novel architectures such as compute-optimized graphics processors, many-core processor/accelerators and, even special-purpose function units, the vast majority of scientific calculations are still performed on high-performance, commodity server processors. Even in the cases of applications which have been ported to new architectures, frequent serial sections still require strong server-class processor cores to compute as fast as possible. In this paper we report on a set of benchmark studies which evaluate Intel's latest Skylake Xeon server processor. Skylake represents a significant change in the Xeon product line with wider SIMD vector units, a redesigned cache architecture, and, an increased number of memory channels. The wider vector units provide 2x improvement for some compute-intensive applications and the combined memory changes can provide close to 2x the memory bandwidth. We evaluate these new hardware features on several HPC-relevant mini-applications and benchmarks, including, STREAM, LULESH, XSBench, HPCG and SW4Lite. Together, the new hardware functions provide up to 1.8x speedup on HPC benchmark codes when compared with the previous generation Haswell processor core, providing much greater utility to a broader range of HPC applications that rely on this class of compute node.