Implementation and Numerical Techniques for One EFlop/s HPL-AI Benchmark on Fugaku

Implementation and Numerical Techniques for One EFlop/s HPL-AI Benchmark on Fugaku
复制标题

Fugaku 上 EFlop/s HPL-AI 基准的实现和数值技术

DOI:
10.1109/scala51936.2020.00014
复制
发表时间:
2020
期刊:
2020 IEEE/ACM 11th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems (ScalA)
影响因子:
--
通讯作者:
Toshiyuki Imamura
Toshiyuki Imamura
中科院分区:
--
文献类型:
--
作者:
Shuhei Kudo;Keigo Nitadori;Takuya Ina;Toshiyuki Imamura

文献摘要

参考文献

被引文献

相似文献

我们在超级计算机Fugaku上的HPL-AI的性能基准获得了第55位前500名。有效的性能为1.42 Eflop/s,并且在浮点算术基准中超过Exa规模的墙壁的首要成就。由于HPL-AI是全新的,并且没有针对大型系统的参考代码,因此从低精确的数值角度来看,大规模基准中存在一些挑战。仅用FP32或FP16的操作替换FP64操作是不够的。至少,我们需要进行较低精确算术的周到的数值分析,并在诸如Fugaku等广泛计算上引入优化技术。这项研究对Fugaku上EXA规模的基准进行了一些技术分析和有关准确性问题,实施,绩效改善的见解。
Our performance benchmark of HPL-AI on the supercomputer Fugaku was awarded the 55th Top500. The effective performance was 1.42 EFlop/s, and the world's first achievement to exceed the wall of Exa-scale in a floating-point arithmetic benchmark. Because HPL-AI is brand new and has no reference code for large systems, several challenges exists in the large-scale benchmark from a low-precision numerical viewpoint. It is not sufficient to replace FP64 operations solely with those of FP32 or FP16. At the least, we need thoughtful numerical analysis for lower-precision arithmetic and the introduction of optimization techniques on extensive computing such as on Fugaku. This study presents some technical analysis and insights on the accuracy issues, implementation, performance improvement, and report on the Exa-scale benchmark on Fugaku.
DOI: 10.1088/0266-5611/13/2/022
发表时间: 1997
期刊: Inverse Problems
影响因子: 2.1
作者:
通讯作者: --