Porting DDalphaAMG solver to K computer

Porting DDalphaAMG solver to K computer
复制标题

将 DDalphaAMG 求解器移植到 K 计算机

DOI:
--
复制
发表时间:
2018
期刊:
Proceedings of The 36th Annual International Symposium on Lattice Field Theory — PoS(LATTICE2018)
影响因子:
--
通讯作者:
I. Kanamori
I. Kanamori
中科院分区:
--
文献类型:
--
作者:
K. Ishikawa;I. Kanamori

文献摘要

参考文献

被引文献

相似文献

我们将Domain-Decomposed-alpha-AMG求解器移植到K计算机。该系统具有8个核心和每个节点16 GB内存,其中理论峰值为128 GFlops(总共82,944个节点)。它的特点是每个内核多达256个寄存器,最大可达0.5字节/触发器比率,需要与其他机器不同的调优。为了使用更多的寄存器,我们改变了一些数据结构,并使用内部函数重写了矩阵向量运算。对于包括设置在内的十二次求解,性能提高了两倍以上。优化后的效率仍为5%左右,低于K计算机的混合精度求解器的22%。但是,对于物理点配置,吞吐量要高出两倍以上。
We port Domain-Decomposed-alpha-AMG solver to the K computer. The system has 8 cores and 16 GB memory per node, of which theoretical peak is 128 GFlops (82,944 nodes in total). Its feature, as many as 256 registers per core and as large as 0.5 byte/Flop ratio, requires a different tuning from other machines. In order to use more registers, we change some of the data structure and rewrite matrix-vector operations with intrinsics. The performance is improved by more than a factor two for twelve solves including the setup. The efficiency is still about 5% after the optimization, which is lower than a previously tuned mixed precision solver for the K computer, 22%. The throughput is, however, more than two times better for a physical point configuration.
DOI: 10.1137/130919507
发表时间: 2013-03
期刊: SIAM J. Sci. Comput.
影响因子: --
作者:
A. Frommer;K. Kahl;S. Krieg;B. Leder;M. Rottmann
通讯作者: A. Frommer;K. Kahl;S. Krieg;B. Leder;M. Rottmann