Implementation and Numerical Techniques for One EFlop/s HPL-AI Benchmark on Fugaku
Implementation and Numerical Techniques for One EFlop/s HPL-AI Benchmark on Fugaku
复制标题
Fugaku 上 EFlop/s HPL-AI 基准的实现和数值技术
DOI:
10.1109/scala51936.2020.00014
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Toshiyuki Imamura
中科院分区:
文献类型:
--
作者:
Shuhei Kudo;Keigo Nitadori;Takuya Ina;Toshiyuki Imamura
Our performance benchmark of HPL-AI on the supercomputer Fugaku was awarded the 55th Top500. The effective performance was 1.42 EFlop/s, and the world's first achievement to exceed the wall of Exa-scale in a floating-point arithmetic benchmark. Because HPL-AI is brand new and has no reference code for large systems, several challenges exists in the large-scale benchmark from a low-precision numerical viewpoint. It is not sufficient to replace FP64 operations solely with those of FP32 or FP16. At the least, we need thoughtful numerical analysis for lower-precision arithmetic and the introduction of optimization techniques on extensive computing such as on Fugaku. This study presents some technical analysis and insights on the accuracy issues, implementation, performance improvement, and report on the Exa-scale benchmark on Fugaku.
影响因子:
2.1
作者:
通讯作者:
--