Quadruple-precision BLAS using Bailey's arithmetic with FMA instruction: its performance and applications

Quadruple-precision BLAS using Bailey's arithmetic with FMA instruction: its performance and applications
复制标题

使用贝利算法和 FMA 指令的四精度 BLAS:性能和应用

DOI:
10.1109/ipdpsw.2017.42
复制
发表时间:
2017
期刊:
Parallel and Distributed Processing Symposium Workshops (IPDPSW), 2017 IEEE International
影响因子:
--
通讯作者:
Toshiyuki Imamura
Toshiyuki Imamura
中科院分区:
--
文献类型:
--
作者:
Susumu Yamada;Takuya Ina;Narimasa Sasa;Yasuhiro Idomura;Masahiko Machida;Toshiyuki Imamura

文献摘要

相似文献

当在处理器单元上执行浮点运算时,每次计算都会出现舍入和截断错误。这些误差会导致需要大量计算的大型模拟中的精度问题。因此,我们开发了基于贝利双精度算术的四精度基本线性代数子程序(QPBLAS)。贝利算术的乘法运算是通过24次双精度运算实现的。当使用FMA(融合乘加)指令时,我们可以将运算次数减少大约一半。因此,我们使用FMA指令开发QPBLAS并评估其性能。结果表明,使用FMA的QPBLAS基本上比QPBLAS快。此外,当我们使用FMA将四精度特征值求解器QPEigenK的QPBLAS替换为QPBLAS时,我们可以获得大约10~20%的加速。
When a floating-point arithmetic is executed on a processor unit, round-off and truncation errors occur every calculation. These errors cause a precision issue in a large simulation which requires a great number of calculations. Therefore, we have developed the quadruple-precision basic linear algebra subprograms (QPBLAS) based on Bailey's double-double arithmetic. The multiplication operation of Bailey's arithmetic is realized by 24 double-precision operations. When using an FMA (fused multiply-add) instruction, we can reduce the number of operations in about half. Therefore, we develop QPBLAS using the FMA instruction and evaluate its performance. The result shows that the QPBLAS using FMA is basically faster than QPBLAS. Moreover, when we replace QPBLAS of the quadruple-precision eigenvalue solver QPEigenK with QPBLAS using FMA, we can obtain about 10~20% speedup.