On-FPGA Training with Ultra Memory Reduction: A Low-Precision Tensor Method

On-FPGA Training with Ultra Memory Reduction: A Low-Precision Tensor Method
复制标题

DOI:
--
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Kaiqi Zhang;Cole Hawkins;Xiyuan Zhang;Cong Hao;Zheng Zhang
Kaiqi Zhang;Cole Hawkins;Xiyuan Zhang;Cong Hao;Zheng Zhang
中科院分区:
其他
文献类型:
--
作者:
Kaiqi Zhang;Cole Hawkins;Xiyuan Zhang;Cong Hao;Zheng Zhang

文献摘要

相似文献

为了在边缘设备上对神经网络进行节能和实时推理,已经开发了各种硬件加速器。然而,大多数训练都是在高性能的GPU或服务器上进行的,巨大的内存和计算成本阻碍了在边缘设备上训练神经网络。本文提出了一种新的基于张量的训练框架,该框架在训练过程中提供了数量级的内存减少。我们提出了一种新的秩自适应张量化神经网络模型,并设计了一种硬件友好的低精度算法来训练该模型。我们提供了一个现场可编程门阵列加速器来演示这种训练方法在边缘设备上的好处。与嵌入式CPU相比,我们的初步FPGA实现获得了59倍的加速比和123倍的能耗,与标准的全尺寸训练相比,内存减少了292倍。
Various hardware accelerators have been developed for energy-efficient and real-time inference of neural networks on edge devices. However, most training is done on high-performance GPUs or servers, and the huge memory and computing costs prevent training neural networks on edge devices. This paper proposes a novel tensor-based training framework, which offers orders-of-magnitude memory reduction in the training process. We propose a novel rank-adaptive tensorized neural network model, and design a hardware-friendly low-precision algorithm to train this model. We present an FPGA accelerator to demonstrate the benefits of this training method on edge devices. Our preliminary FPGA implementation achieves $59\times$ speedup and $123\times$ energy reduction compared to embedded CPU, and $292\times$ memory reduction over a standard full-size training.