Accurate and Efficient Floating Point Summation

Accurate and Efficient Floating Point Summation
复制标题

DOI:
10.1137/s1064827502407627
复制
发表时间:
2003-04
期刊:
SIAM J. Sci. Comput.
影响因子:
--
通讯作者:
J. Demmel;Yozo Hida
J. Demmel;Yozo Hida
中科院分区:
其他
文献类型:
--
作者:
J. Demmel;Yozo Hida

文献摘要

被引文献

相似文献

我们提出并分析了几个简单的算法,准确地计算n个浮点数的总和使用更广泛的累加器。令f和F分别为被加数和累加器中的有效位数。然后假设逐渐下溢,没有溢出,和四舍五入到最近的算术,高达约2F-f数可以通过简单地按指数的降序求和来精确地相加,产生在最后一位(ulps)的大约1.5个单位内正确的和。我们将此结果应用到IEEE浮点标准中的浮点格式。例如,使用双精度和排序计算的长度最多为33的单精度向量的点积保证正确率接近1.5 ulps。如果使用双倍扩展精度,则向量长度可以高达65,537。我们还研究了如何在保持准确性的同时减少或消除排序成本。
We present and analyze several simple algorithms for accurately computing the sum of n floating point numbers using a wider accumulator. Let f and F be the number of significant bits in the summands and the accumulator, respectively. Then assuming gradual underflow, no overflow, and round-to-nearest arithmetic, up to approximately 2F-f numbers can be added accurately by simply summing the terms in decreasing order of exponents, yielding a sum correct to within about 1.5 units in the last place (ulps). We apply this result to the floating point formats in the IEEE floating point standard. For example, a dot product of single precision vectors of length at most 33 computed using double precision and sorting is guaranteed correct to nearly 1.5 ulps. If double-extended precision is used, the vector length can be as large as 65,537. We also investigate how the cost of sorting can be reduced or eliminated while retaining accuracy.