Proposal of “effective floating-point number” for approximate algebraic computation
Proposal of “effective floating-point number” for approximate algebraic computation
复制标题
近似代数计算“有效浮点数”的提议
DOI:
10.1145/271130.271136
复制
发表时间:
1997
期刊:
影响因子:
--
通讯作者:
Tateaki Sasaki
中科院分区:
文献类型:
--
作者:
F. Kako;Tateaki Sasaki
The floating-point number, or float number in short, is very useful but it suffers two kinds of errors, round-off error and cancellation error. In approximate algebraic computation using float numbers [3, 4], the cancellation error is very dangerous but the round-off error is not, and monitoring the occurrence and the amount of cancellation error during the computation is crucially important. The most famous numeric arithmetic that was expected to solve the error problems of float numbers is interval arithmetic [2, 1]. The interval arithmetic gives us a rigorous upper bound of error and it seems to be complete from the viewpoint of mathematics. It is, however, not so useful because the u p p e r bound obtained is usually too large overestimation. A less famous one is "error number" arithmetic which has been implemented in SMP [4], the ancestor of Mathematica. The error number is considered to represent floating-point numbers distributed statistically around f with the variance a, and it is represented as (f, a). For convenience, we call the error number variance floating-point number, or vfloat number in short. The effective floating-point number, or efloat number in short, which we propose is for solving the problem of cancellation error, but it is useless for round-off error. The round-off error is caused by rounding the tail bit of bit sequence representing the mantissa, while the cancellation error is due to the cancellation of leading bits of the bit sequence. We count the number of leading bits canceled in the addition and subtraction and reserve it in the representation of float number, such as (f,n), where f is the conventional float number and n is the number of canceled bits. In the computation of efloat number, we consider that only the leading M-n-1 bits are effective, where M is the length of the bit sequence representing the mantissa. Counting the number of canceled bits is easy by hardware but not so easy by software. In software implementation of efloat numbers, we represent an efloat number as (f, e), where e '~ 2n-Mill hence e represents the cancellation error approximately. We set e initially as e = 2-Mill. We denote a real interval [a,b] = [ c-w,c + w], a real vfloat number (f , a) and a real efloat number (f,e) as [c, W]A, (f, a)A and (f, e)A, respectively, and call absolute width/variance/error representations. Let r = w/Ic[, 7" …