Approximate Constant-Coefficient Multiplication Using Hybrid Binary-Unary Computing for FPGAs

Approximate Constant-Coefficient Multiplication Using Hybrid Binary-Unary Computing for FPGAs
复制标题

使用 FPGA 混合二元-一元计算近似常数系数乘法

DOI:
--
复制
发表时间:
2021
影响因子:
2.3
通讯作者:
K. Bazargan
K. Bazargan
中科院分区:
计算机科学3区
文献类型:
--
作者:
S. R. Faraji;Pierre Abillama;K. Bazargan

文献摘要

被引文献

相似文献

乘法器用于几乎所有的数字信号处理(DSP)应用,如图像和视频处理。乘法器效率直接影响此类应用的整体性能,特别是在需要实时处理的情况下(如4K视频处理),或者在硬件资源有限的情况下(如移动的和物联网设备)。我们提出了一种新颖的,低成本,低能耗,高速近似常数系数乘法器(CCM)使用混合二进制一元编码方法。所提出的方法使用简单的路由网络实现CCM,在一元域中没有逻辑门,这导致比Xilinx LogiCORE IP CCM和基于表的KCM CCM(Flopoco)平均更有效的乘法器。我们评估所提出的乘法器上的2-D离散余弦变换算法作为一个共同的DSP模块。布线后FPGA测试结果表明,该乘法器能使二维离散余弦变换的{面积、面积×时延、功耗和能量-时延积}平均提高{30%,33%,30%,31%}。此外,所提出的2-D离散余弦变换的吞吐量是平均5%以上,使用基于表的KCM CCM实现的二进制架构。我们将表明,我们的方法相比,二进制实现时,实现DCT核心的可路由性问题较少。
Multipliers are used in virtually all Digital Signal Processing (DSP) applications such as image and video processing. Multiplier efficiency has a direct impact on the overall performance of such applications, especially when real-time processing is needed, as in 4K video processing, or where hardware resources are limited, as in mobile and IoT devices. We propose a novel, low-cost, low energy, and high-speed approximate constant coefficient multiplier (CCM) using a hybrid binary-unary encoding method. The proposed method implements a CCM using simple routing networks with no logic gates in the unary domain, which results in more efficient multipliers compared to Xilinx LogiCORE IP CCMs and table-based KCM CCMs (Flopoco) on average. We evaluate the proposed multipliers on 2-D discrete cosine transform algorithm as a common DSP module. Post-routing FPGA results show that the proposed multipliers can improve the {area, area × delay, power consumption, and energy-delay product} of a 2-D discrete cosine transform on average by {30%, 33%, 30%, 31%}. Moreover, the throughput of the proposed 2-D discrete cosine transform is on average 5% more than that of the binary architecture implemented using table-based KCM CCMs. We will show that our method has fewer routability issues compared to binary implementations when implementing a DCT core.