Efficient Reproducible Floating Point Summation and BLAS

Efficient Reproducible Floating Point Summation and BLAS
复制标题

高效、可重复的浮点求和和 BLAS

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
J. Demmel
J. Demmel
中科院分区:
--
文献类型:
--
作者:
Peter Ahrens;Hong Diep Nguyen;J. Demmel

文献摘要

被引文献

相似文献

我们定义了重现性,意味着从同一程序的多个运行中获得相同的结果,也许使用不同的硬件资源或其他更改,这些更改应该不会更改答案并行计算资源的浮动点增加,使获得可重复性成为挑战数字的向量,或更复杂的操作,例如基本的线性代数子程序(BLA)。点标准754-2008。需要一个只有6个浮点字表示的“累加器”(如果需要更高的精度,则可以使用6个字的累加器。最终错误的限制可能比传统求和的误差要小10-8倍。 (毛布拉斯)和性能结果。 CPU在3.4 GHz和256 kb L2缓存下运行。
We define reproducibility to mean getting bitwise identical results from multiple runs of the same program, perhaps with different hardware resources or other changes that should ideally not change the answer. Many users depend on reproducibility for debugging or correctness [1]. However, dynamic scheduling of parallel computing resources, combined with nonassociativity of floating point addition, makes attaining reproducibility a challenge even for simple operations like summing a vector of numbers, or more complicated operations like the Basic Linear Algebra Subprograms (BLAS). We describe an algorithm that computes a reproducible sum of floating point numbers, independent of the order of summation. The algorithm depends only on a subset of the IEEE Floating Point Standard 754-2008. It is communication-optimal, in the sense that it does just one pass over the data in the sequential case, or one reduction operation in the parallel case, requiring an “accumulator” represented by just 6 floating point words (more can be used if higher precision is desired). The arithmetic cost with a 6-word accumulator is 7n floating point additions to sum n words, and (in IEEE double precision) the final error bound can be up to 10−8 times smaller than the error bound for conventional summation. We describe the basic summation algorithm, the software infrastructure used to build reproducible BLAS (ReproBLAS), and performance results. For example, when computing the dot product of 4096 double precision floating point numbers, we get an 4x slowdown compared to Intel R ©Math Kernel Library (MKL) running on an Intel R ©Core i7-2600 CPU operating at 3.4 GHz and 256 KB L2 Cache.