Design-Space Exploration of Quantized Transposed Convolutional Neural Networks for FPGA-based Systems-on-Chip

Design-Space Exploration of Quantized Transposed Convolutional Neural Networks for FPGA-based Systems-on-Chip
复制标题

DOI:
10.1109/dasc/picom/cbdcom/cy55231.2022.9927825
复制
发表时间:
2022-09
期刊:
2022 IEEE Intl Conf on Dependable, Autonomic and Secure Computing, Intl Conf on Pervasive Intelligence and Computing, Intl Conf on Cloud and Big Data Computing, Intl Conf on Cyber Science and Technology Congress (DASC/PiCom/CBDCom/CyberSciTech)
影响因子:
--
通讯作者:
Cristian Sestito;S. Perri;Rob Stewart
Cristian Sestito;S. Perri;Rob Stewart
中科院分区:
其他
文献类型:
--
作者:
Cristian Sestito;S. Perri;Rob Stewart

文献摘要

相似文献

随着深度学习应用向边缘计算设备的转变,压缩技术已经被引入,以最大限度地减少硬件使用、功耗和延迟。例如,量化使用低数值精度来表示输入、参数和激活函数。转置卷积(TCONVs)为神经网络提供图像上采样能力。然而,TCONV层的精度和性能权衡还没有得到充分的研究,现有的工作评估精度低至8位,但不低于8位。本研究系统地评估了在基于fpga的片上系统(SoC)架构中使用TCONVs实现两层量化解码器时非常低精度的影响。我们评估了量化对吞吐量性能和硬件成本的影响,以及使用相同指标并行化TCONV层计算的影响。结果表明,当处理4位数据时,在Xilinx Zynq-7020 SoC上实现的电路仅使用~15%的逻辑和~7.5%的片上存储器,而相对于8位对应的电路,精度损失可忽略的~2.5%。此外,当输入以4倍的并行度处理时,可以观察到3.5倍的加速。
With the shift of deep learning applications to Edge Computing devices, compression techniques have been introduced to minimize hardware use, power consumption and latency. For example, quantization uses low numeric precision to represent inputs, parameters and activation functions. Transposed Convolutions (TCONVs) provide neural networks with image up-sampling capabilities. However, the accuracy and performance trade-off of TCONV Layers is under-explored, with existing works evaluating down to 8-bit precision but not less. This research systematically evaluates the impact of very low precision when a two-layers quantized decoder, using TCONVs, is implemented within an FPGA-based System-on-Chip (SoC) architecture. We evaluate the quantization impact on throughput performance and hardware costs, as well as the impact of parallelizing the computations of TCONV Layers using the same metrics. Results show that, when 4-bit data are processed, the circuit implemented on a Xilinx Zynq-7020 SoC only uses ~15% of logic and ~7.5% of on-chip memories, at the expense of a negligible ~2.5% accuracy loss with respect to the 8-bit counterpart. Furthermore, 3.5× speed-up is observed when inputs are processed with 4× parallelism.