Accelerating Pipelined Integer and Floating-Point Accumulations in Configurable Hardware with Delayed Addition Techniques

Accelerating Pipelined Integer and Floating-Point Accumulations in Configurable Hardware with Delayed Addition Techniques
复制标题

DOI:
10.1109/12.841125
复制
发表时间:
2000-03
期刊:
IEEE Trans. Computers
影响因子:
--
通讯作者:
Zhen Luo;M. Martonosi
Zhen Luo;M. Martonosi
中科院分区:
其他
文献类型:
--
作者:
Zhen Luo;M. Martonosi

文献摘要

被引文献

相似文献

可配置硬件中的算术计算的速度受到进位传播的限制,即使在最近的FPGA中发现了专用硬件。本文提出并评估了一种称为延迟加法的方法,该方法减少了进位传播瓶颈,提高了算术计算的性能。我们的方法采用了华莱士树的思想,以中间形式存储结果,并延迟加法,直到一个重复的计算,如积累或点积结束,这有效地消除了从计算的关键path. We提出的整数和浮点设计,使用我们的技术进行传播开销。我们的流水线整数乘累加(MAC)的设计是基于一个相当传统的乘法器的设计,但延迟以及。该设计在XC 4036 xla-9 FPGA上实现了72 MHz的时钟速率,在XV 300 epq 240 -8 FPGA上实现了170 MHz的时钟速率。接下来,我们提出了一个基于延迟加法的32位浮点累加器。这里,延迟加法需要一种新的对齐技术,该技术将输入的操作数从累加结果中重新排列。该设计的保守版本在XC 4036 xla-9 FPGA上实现40 MHz时钟速率,在XV 100 epq 240 -8 FPGA上实现97 MHz时钟速率。我们还提出了一个32位浮点累加器的设计与编译器管理的溢出避免,实现了80 MHz的时钟频率上的XC 4036 xla-9 FPGA和150 MHz的时钟频率上的XCV 100 epq 240 -8 FPGA。
The speed of arithmetic calculations in configurable hardware is limited by carry propagation, even with the dedicated hardware found in recent FPGAs. This paper proposes and evaluates an approach called delayed addition that reduces the carry-propagation bottleneck and improves the performance of arithmetic calculations. Our approach employs the idea used in Wallace trees to store the results in an intermediate form and delay addition until the end of a repeated calculation such as accumulation or dot-product; this effectively removes carry propagation overhead from the calculation's critical path. We present both integer and floating-point designs that use our technique. Our pipelined integer multiply-accumulate (MAC) design is based on a fairly traditional multiplier design, but with delayed addition as well. This design achieves a 72 MHz clock rate on an XC4036xla-9 FPGA and 170 MHz clock rate on an XV300epq240-8 FPGA. Next, we present a 32-bit floating-point accumulator based on delayed addition. Here, delayed addition requires a novel alignment technique that decouples the incoming operands from the accumulated result. A conservative version of this design achieves a 40 MHz clock rate on an XC4036xla-9 FPGA and 97 MHz clock rate on an XV100epq240-8 FPGA. We also present a 32-bit floating-point accumulator design with compiler-managed overflow avoidance that achieves a 80 MHz clock rate on an XC4036xla-9 FPGA and 150 MHz clock rate on an XCV100epq240-8 FPGA.