Reproducible, Accurately Rounded and Efficient BLAS

Reproducible, Accurately Rounded and Efficient BLAS
复制标题

可重复、精确舍入且高效的 BLAS

DOI:
--
复制
发表时间:
2016
期刊:
Euro-Par Workshops
影响因子:
--
通讯作者:
David Parello
David Parello
中科院分区:
--
文献类型:
--
作者:
Chemseddine Chohra;P. Langlois;David Parello

文献摘要

被引文献

相似文献

在平行计算中,数值可重复性失败会增加,因为浮点求和是非缔合性的。大规模平行和优化的执行会动态修改浮点操作顺序。因此,数值结果可能会从一个运行变为另一种运行。我们建议通过尽可能扩展IEEE-754正确的圆形属性到更大的操作序列来确保可重复性。我们介绍了稀有的 - 可重现的,准确的,圆形的,有效的布拉),从最近的准确有效的求和算法中受益。提出了1级(ASUM,DOT和NRM2)和2级(GEMV)例程的解决方案。与英特尔MKL库和其他现有可重复的算法相比,研究了它们的性能。对于共享和分布式存储器并行系统,我们在最坏的情况下表现出2(IME)的额外成本为2(IME),这对于广泛的应用程序令人满意。对于Intel Xeon Phi加速器,观察到更大的外部成本(4(IME)至6(IMES)),至少对调试和验证步骤至少有帮助。
Numerical reproducibility failures rise in parallel computation because floating-point summation is non-associative. Massively parallel and optimized executions dynamically modify the floating-point operation order. Hence, numerical results may change from one run to another. We propose to ensure reproducibility by extending as far as possible the IEEE-754 correct rounding property to larger operation sequences. We introduce our RARE-BLAS (Reproducible, Accurately Rounded and Efficient BLAS) that benefits from recent accurate and efficient summation algorithms. Solutions for level 1 (asum, dot and nrm2) and level 2 (gemv) routines are presented. Their performance is studied compared to the Intel MKL library and other existing reproducible algorithms. For both shared and distributed memory parallel systems, we exhibit an extra-cost of 2( imes ) in the worst case scenario, which is satisfying for a wide range of applications. For Intel Xeon Phi accelerator a larger extra-cost (4( imes ) to 6( imes )) is observed, which is still helpful at least for debugging and validation steps.