QPA: A Quantization-Aware Piecewise Polynomial Approximation Methodology for Hardware-Efficient Implementations

QPA: A Quantization-Aware Piecewise Polynomial Approximation Methodology for Hardware-Efficient Implementations
复制标题

QPA:一种用于硬件高效实现的量化感知分段多项式逼近方法

DOI:
10.1109/tvlsi.2023.3277023
复制
发表时间:
2023
影响因子:
2.8
通讯作者:
Li Du
Li Du
中科院分区:
工程技术2区
文献类型:
--
作者:
Haoran Geng;Xiaoliang Chen;Ning Zhao;Yuan Du;Li Du

文献摘要

相似文献

非线性函数的分段多项式逼近在高精度计算中起着重要作用。在这篇文章中,我们提出了QPA,一个错误平坦量化感知PPA方法的集成,以生成针对任何多项式阶数的高效硬件实现的优化系数。QPA结合了四个关键特征来最小化拟合误差和硬件成本,包括使用Remez算法来计算极小极大拟合多项式,将拟合和量化操作相结合以获得误差平坦特性,为每个乘法器分配特定的系数位宽以降低硬件成本,以及微调截断系数以进一步降低拟合误差。实验结果表明,我们的方法一致实现最低的拟合误差相比,最先进的误差平坦分段逼近方法。我们综合了建议的设计与28纳米台积电CMOS工艺。结果表明,所提出的设计实现了高达37.0%的面积减少和50.5%的功耗降低相比,最先进的误差平坦分段线性(PWL)的方法,和高达27.0%的面积减少,21.4%的延迟减少,和20.8%的功耗减少相比,最先进的误差平坦分段二次(PWQ)的方法。
Piecewise polynomial approximation (PPA) on nonlinear functions plays an important role in high-precision computing. In this article, we proposed QPA, an integration of error-flattened quantization-aware PPA methods, to generate the optimized coefficients for efficient hardware implementations targeting any polynomial order. QPA incorporated four key features to minimize the fitting error and the hardware cost, including using the Remez algorithm to compute the minimax fitting polynomial, combining the fitting and quantization operations to get an error-flattened characteristic, assigning specific coefficient bit width to each multiplier to reduce the hardware cost, and fine-tuning the truncated coefficients to further reduce the fitting error. Experimental results showed that our methods consistently achieved the lowest fitting error compared with the state-of-the-art error-flattened piecewise approximation methods. We synthesized the proposed designs with 28-nm TSMC CMOS technology. The results showed that the proposed designs achieved up to 37.0% area reduction and 50.5% power consumption reduction compared to the state-of-the-art error-flattened piecewise linear (PWL) method, and up to 27.0% area reduction, 21.4% delay reduction, and 20.8% power consumption reduction compared to the state-of-the-art error-flattened piecewise quadratic (PWQ) method.