FPGA Implementation for the Sigmoid with Piecewise Linear Fitting Method Based on Curvature Analysis

FPGA Implementation for the Sigmoid with Piecewise Linear Fitting Method Based on Curvature Analysis
复制标题

基于曲率分析的分段线性拟合Sigmoid函数的FPGA实现

DOI:
10.3390/electronics11091365
复制
发表时间:
2022
期刊:
影响因子:
2.9
通讯作者:
Qinglin Wang
Qinglin Wang
中科院分区:
工程技术3区
文献类型:
--
作者:
Zerun Li;Yang Zhang;Bingcai Sui;Zuocheng Xing;Qinglin Wang

文献摘要

相似文献

Sigmoid激活功能在神经网络中很受欢迎,但其复杂性限制了硬件的实现和速度。在本文中,我们使用曲率值将Sigmoid函数分为不同的段,并采用最小二乘方法来求解每个段中分段线性拟合函数的表达式。然后,我们采用具有最大绝对错误和平均绝对错误的优化方法,以选择具有指定数量段数的适当函数表达式。最后,我们在现场可编程栅极阵列(FPGA)开发平台上实现Sigmoid函数,并同时应用算术(乘法和添加)和范围选择的并行操作。 FPGA实现结果表明,我们设计的时钟频率高达208.3 MHz,而端到端延迟仅为9.6 ns。我们基于曲率分析(PWLC)的分段线性拟合方法在97.51%的MNIST数据集上具有深度神经网络(DNN)(DNN)和98.65%的识别精度,并具有卷积神经网络(CNN)的98.65%。实验结果表明,我们的Sigmoid功能的FPGA设计可以获得最低的延迟,减少绝对错误并获得高识别精度,而硬件成本在实际应用中可以接受。
The sigmoid activation function is popular in neural networks, but its complexity limits the hardware implementation and speed. In this paper, we use curvature values to divide the sigmoid function into different segments and employ the least squares method to solve the expressions of the piecewise linear fitting function in each segment. We then adopt an optimization method with maximum absolute errors and average absolute errors to select an appropriate function expression with a specified number of segments. Finally, we implement the sigmoid function on the field-programmable gate array (FPGA) development platform and apply parallel operations of arithmetic (multiplying and adding) and range selection at the same time. The FPGA implementation results show that the clock frequency of our design is up to 208.3 MHz, while the end-to-end latency is just 9.6 ns. Our piecewise linear fitting method based on curvature analysis (PWLC) achieves recognition accuracy on the MNIST dataset of 97.51% with a deep neural network (DNN) and 98.65% with a convolutional neural network (CNN). Experimental results demonstrate that our FPGA design of sigmoid function can obtain the lowest latency, reduce absolute errors, and achieve high recognition accuracies, while the hardware cost is acceptable in practical applications.