Reduced-Area Constant-Coefficient and Multiple-Constant Multipliers for Xilinx FPGAs with 6-Input LUTs

Reduced-Area Constant-Coefficient and Multiple-Constant Multipliers for Xilinx FPGAs with 6-Input LUTs
复制标题

适用于具有 6 输入 LUT 的 Xilinx FPGA 的面积减小常数系数和多重常数乘法器

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
E. G. Walters
E. G. Walters
中科院分区:
--
文献类型:
--
作者:
E. G. Walters

文献摘要

被引文献

相似文献

乘以一个常数是在现场可编程门阵列(fpga)中实现的许多信号、图像和视频处理应用程序的常见操作。常系数乘法器(kcm)通常在逻辑结构中使用查找表(lut)实现,为通用乘法保留嵌入式硬乘法器。本文描述了以前工作中的一个双操作数加法电路,并展示了如何使用它来生成和添加预先计算的部分乘积来实现kcm。提出了一种新的带负常数的kcm偏积预计算方法。然后将这些kcm扩展为具有2到8个系数,这些系数可以在运行时由控制信号选择,以实现时间复用的多常数乘法。综合结果表明,所提出的流水线KCMs平均使用的lut减少了27.4%,并且具有比LogiCORE IP KCMs低12%的中位数lut延迟产品。与基于LogiCORE IP的最佳替代方案相比,具有2到8个可选择系数的流水线KCMs使用的lut减少了46%到70%,并且大多数比使用带有系数查找功能的LogiCORE IP乘法器更快。在相同的操作数大小和可选择系数数量下,它们的切片比最小的流水线加法图(PAG)融合设计少22%到57%,运行速度比最快的PAG融合设计快7%到30%。对于kcm和具有给定操作数大小的可选择系数的kcm,对于所有正负常数值,lut的放置和路由保持相同,这有利于运行时部分重新配置。
Multiplication by a constant is a common operation for many signal, image, and video processing applications that are implemented in field-programmable gate arrays (FPGAs). Constant-coefficient multipliers (KCMs) are often implemented in the logic fabric using lookup tables (LUTs), reserving embedded hard multipliers for general-purpose multiplication. This paper describes a two-operand addition circuit from previous work and shows how it can be used to generate and add pre-computed partial products to implement KCMs. A novel method for pre-computing partial products for KCMs with a negative constant is also presented. These KCMs are then extended to have two to eight coefficients that may be selected by a control signal at runtime to implement time-multiplexed multiple-constant multiplication. Synthesis results show that proposed pipelined KCMs use 27.4% fewer LUTs on average and have a median LUT-delay product that is 12% lower than comparable LogiCORE IP KCMs. Proposed pipelined KCMs with two to eight selectable coefficients use 46% to 70% fewer LUTs than the best LogiCORE IP based alternative and most are faster than using a LogiCORE IP multiplier with a coefficient lookup function. They also outperform the state-of-the-art in the literature, using 22% to 57% fewer slices than the smallest pipelined adder graph (PAG) fusion designs and operate 7% to 30% faster than the fastest PAG fusion designs for the same operand size and number of selectable coefficients. For KCMs and KCMs with selectable coefficients of a given operand size, the placement and routing of LUTs remains the same for all positive and negative constant values, which is advantageous for runtime partial reconfiguration.