Development of Element-by-Element Kernel Algorithms in Unstructured Implicit Low-Order Finite-Element Earthquake Simulation for Many-Core Wide-SIMD CPUs

Development of Element-by-Element Kernel Algorithms in Unstructured Implicit Low-Order Finite-Element Earthquake Simulation for Many-Core Wide-SIMD CPUs
复制标题

DOI:
10.1007/978-3-030-22734-0_20
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
K. Fujita;Masashi Horikoshi;T. Ichimura;L. Meadows;K. Nakajima;M. Hori;Lalith Maddegedara
K. Fujita;Masashi Horikoshi;T. Ichimura;L. Meadows;K. Nakajima;M. Hori;Lalith Maddegedara
中科院分区:
其他
文献类型:
--
作者:
K. Fujita;Masashi Horikoshi;T. Ichimura;L. Meadows;K. Nakajima;M. Hori;Lalith Maddegedara

文献摘要

相似文献

矩阵向量积中的逐元素(EBE)内核的加速对于非结构化隐式有限元应用的高性能是必不可少的。然而,EBE内核并不直接获得高性能,由于随机数据访问与数据递归。在本文中,我们开发的方法来规避这些数据竞争的高性能的多核CPU架构与广泛的SIMD单元。开发的EBE内核在基于Intel Xeon Phi Knights Landing的Oakforest-PACS和基于Intel Skylake Xeon Gold处理器的系统上分别达到FP 32峰值的16.3%和20.9%。这导致了2.88倍的加速比基线内核和2.03倍的加速比的整个有限元应用程序Oakforest-PACS。城市地震模拟使用开发的有限元应用程序的一个例子。
Acceleration of the Element-by-Element (EBE) kernel in matrix-vector products is essential for high-performance in unstructured implicit finite-element applications. However, the EBE kernel is not straightforward to attain high performance due to random data access with data recurrence. In this paper, we develop methods to circumvent these data races for high performance on many-core CPU architectures with wide SIMD units. The developed EBE kernel attains 16.3% and 20.9% of FP32 peak on Intel Xeon Phi Knights Landing based Oakforest-PACS and Intel Skylake Xeon Gold processor based system, respectively. This leads to 2.88-fold speedup over the baseline kernel and 2.03-fold speedup of the whole finite-element application on Oakforest-PACS. An example of urban earthquake simulation using the developed finite-element application is shown.