Proposal of “effective floating-point number” for approximate algebraic computation

Proposal of “effective floating-point number” for approximate algebraic computation
复制标题

近似代数计算“有效浮点数”的提议

DOI:
10.1145/271130.271136
复制
发表时间:
1997
期刊:
SIGSAM Bull.
影响因子:
--
通讯作者:
Tateaki Sasaki
Tateaki Sasaki
中科院分区:
--
文献类型:
--
作者:
F. Kako;Tateaki Sasaki

文献摘要

被引文献

相似文献

浮点数,简称浮点数,是非常有用的,但它遭受两种错误,舍入错误和取消错误。在使用浮点数[3,4]的近似代数计算中,消除误差是非常危险的,但舍入误差不是,并且在计算期间监视消除误差的发生和量是至关重要的。最著名的数值算法,预计将解决错误问题的浮点数是区间算术[2,1]。区间算术给出了严格的误差上界,从数学的观点看,它似乎是完备的。然而,它不是那么有用,因为得到的上界通常是太大的高估。一个不太有名的是“错误数”算法,它已经在Mathematica的祖先SMP [4]中实现。错误数被认为是表示以方差a统计地分布在f周围的浮点数,并且它被表示为(f,a)。为了方便起见,我们将错误数称为方差浮点数,简称为vfloat数。我们提出的有效浮点数,简称浮点数,是为了解决抵消误差的问题,但它对舍入误差是无用的。舍入误差是由表示尾数的比特序列的尾比特舍入引起的,而消除误差是由于比特序列的前导比特的消除。我们计算在加法和减法中取消的前导位的数量,并将其保留在浮点数的表示中,例如(f,n),其中f是传统的浮点数,n是取消的位的数量。在浮点数的计算中,我们认为只有前M-n-1位是有效的,其中M是表示尾数的位序列的长度。计算取消位的数量对于硬件来说很容易,但对于软件来说就不那么容易了。在浮式数的软件实现中,我们将浮式数表示为(f,e),其中e '~ 2n-Mill,因此e近似表示抵消误差。我们初始设置e为e = 2-Mill。我们将一个真实的区间[a,B] = [ c-w,c + w],一个真实的vfloat数(f,a)和一个真实的efloat数(f,e)分别记为[c,W]A,(f,a)A和(f,e)A,并称之为绝对宽度/方差/误差表示。令r = w/Ic[,7”...
The floating-point number, or float number in short, is very useful but it suffers two kinds of errors, round-off error and cancellation error. In approximate algebraic computation using float numbers [3, 4], the cancellation error is very dangerous but the round-off error is not, and monitoring the occurrence and the amount of cancellation error during the computation is crucially important. The most famous numeric arithmetic that was expected to solve the error problems of float numbers is interval arithmetic [2, 1]. The interval arithmetic gives us a rigorous upper bound of error and it seems to be complete from the viewpoint of mathematics. It is, however, not so useful because the u p p e r bound obtained is usually too large overestimation. A less famous one is "error number" arithmetic which has been implemented in SMP [4], the ancestor of Mathematica. The error number is considered to represent floating-point numbers distributed statistically around f with the variance a, and it is represented as (f, a). For convenience, we call the error number variance floating-point number, or vfloat number in short. The effective floating-point number, or efloat number in short, which we propose is for solving the problem of cancellation error, but it is useless for round-off error. The round-off error is caused by rounding the tail bit of bit sequence representing the mantissa, while the cancellation error is due to the cancellation of leading bits of the bit sequence. We count the number of leading bits canceled in the addition and subtraction and reserve it in the representation of float number, such as (f,n), where f is the conventional float number and n is the number of canceled bits. In the computation of efloat number, we consider that only the leading M-n-1 bits are effective, where M is the length of the bit sequence representing the mantissa. Counting the number of canceled bits is easy by hardware but not so easy by software. In software implementation of efloat numbers, we represent an efloat number as (f, e), where e '~ 2n-Mill hence e represents the cancellation error approximately. We set e initially as e = 2-Mill. We denote a real interval [a,b] = [ c-w,c + w], a real vfloat number (f , a) and a real efloat number (f,e) as [c, W]A, (f, a)A and (f, e)A, respectively, and call absolute width/variance/error representations. Let r = w/Ic[, 7" …