Wanted: Floating-Point Add Round-off Error instruction
Wanted: Floating-Point Add Round-off Error instruction
复制标题
需要:浮点加舍入误差指令
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
E. J. Riedy
中科院分区:
文献类型:
--
作者:
Marat Dukhan;R. Vuduc;E. J. Riedy
We propose a new instruction (FPADDRE) that computes the round-o error in oating-point addition. We explain how this instruction benets high-precision arithmetic operations in applications where double precision is not sucient. Performance estimates on Intel Haswell, Intel Skylake, and AMD Steamroller processors, as well as Intel Knights Corner co-processor, demonstrate that such an instruction would improve the latency of double-double addition by up to 55% and increase double-double addition throughput by up to 103%, with smaller, but non-negligible benets for doubledouble multiplication. The new instruction delivers up to 2 speedups on three benchmarks that use high-precision oating-point arithmetic: double-double matrix-matrix multiplication, compensated dot product, and polynomial evaluation via the compensated Horner scheme.