dsODENet: Neural ODE and Depthwise Separable Convolution for Domain Adaptation on FPGAs

dsODENet: Neural ODE and Depthwise Separable Convolution for Domain Adaptation on FPGAs
复制标题

DOI:
10.1109/pdp55904.2022.00031
复制
发表时间:
2022-03
期刊:
2022 30th Euromicro International Conference on Parallel, Distributed and Network-based Processing (PDP)
影响因子:
--
通讯作者:
Hiroki Kawakami;Hirohisa Watanabe;K. Sugiura;Hiroki Matsutani
Hiroki Kawakami;Hirohisa Watanabe;K. Sugiura;Hiroki Matsutani
中科院分区:
其他
文献类型:
--
作者:
Hiroki Kawakami;Hirohisa Watanabe;K. Sugiura;Hiroki Matsutani

文献摘要

相似文献

在边缘环境中,高性能深神经网络(DNN)系统的需求量很高。由于其较高的计算复杂性,在严格限制计算资源的边缘设备上部署DNN是一项挑战。在本文中,我们通过结合最近传播的参数还原技术来得出一个紧凑的DNN模型,称为dsodeNet,称为DSODENET:神经ODE(普通微分方程)和DSC(深度可分离的卷积)。 Neural Ode利用了Resnet和Ode之间的相似性,并且在多层之间共享重量参数的大部分,这大大降低了内存消耗。我们将dsodenet应用于域适应性,作为图像分类数据集的实际用例。我们还为DSodeNet提出了一种基于资源的FPGA设计,其中所有参数和特征地图除了预处理和后处理层,都可以映射到OnChip记忆中。它是在Xilinx ZCU104板上实施的,并根据域的适应准确性,训练速度,FPGA资源利用率和与软件对应物相比进行了评估。结果表明,与我们的基线神经ODE实施相比,DSODENET获得了可比较或稍好的域适应精度,而没有预处理和后处理层的总参数大小降低了54.2%至79.8%。我们的FPGA实施将推理速度加速27.9倍。
High-performance deep neural network (DNN)-based systems are in high demand in edge environments. Due to its high computational complexity, it is challenging to deploy DNNs on edge devices with strict limitations on computational resources. In this paper, we derive a compact while highly-accurate DNN model, termed dsODENet, by combining recently-proposed parameter reduction techniques: Neural ODE (Ordinary Differential Equation) and DSC (Depthwise Separable Convolution). Neural ODE exploits a similarity between ResNet and ODE, and shares most of weight parameters among multiple layers, which greatly reduces the memory consumption. We apply dsODENet to a domain adaptation as a practical use case with image classification datasets. We also propose a resource-efficient FPGA-based design for dsODENet, where all the parameters and feature maps except for pre- and post-processing layers can be mapped onto onchip memories. It is implemented on Xilinx ZCU104 board and evaluated in terms of domain adaptation accuracy, training speed, FPGA resource utilization, and speedup rate compared to a software counterpart. The results demonstrate that dsODENet achieves comparable or slightly better domain adaptation accuracy compared to our baseline Neural ODE implementation, while the total parameter size without pre- and post-processing layers is reduced by 54.2% to 79.8%. Our FPGA implementation accelerates the inference speed by 27.9 times.