Efficient On-Device Training via Gradient Filtering

Efficient On-Device Training via Gradient Filtering
复制标题

DOI:
10.1109/cvpr52729.2023.00371
复制
发表时间:
2023-01
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yuedong Yang-;Guihong Li;R. Marculescu
Yuedong Yang-;Guihong Li;R. Marculescu
中科院分区:
其他
文献类型:
--
作者:
Yuedong Yang-;Guihong Li;R. Marculescu

文献摘要

被引文献

相似文献

尽管它在联邦学习、持续学习和许多其他应用中很重要,但设备上的培训对EdgeAI来说仍然是一个悬而未决的问题。这个问题源于反向传播算法在训练过程中需要大量的操作(例如,浮点乘法和加法)和内存消耗。因此,在本文中,我们提出了一种新的梯度滤波方法,可以实现设备上的CNN模型训练。更准确地说,我们的方法创建了一个特殊的结构,在梯度映射中具有更少的唯一元素,从而显着降低了训练过程中反向传播的计算复杂度和内存消耗。使用多个CNN模型(例如,MobileNet, DeepLabV3, UPerNet)和设备(例如,Raspberry Pi和Jetson Nano)对图像分类和语义分割进行了大量实验,证明了我们的方法的有效性和广泛适用性。例如,与SOTA相比,我们在ImageNet分类上实现了高达19倍的加速和77.1%的内存节省,只有0.1%的准确性损失。最后,我们的方法易于实现和部署;与NVIDIA Jetson Nano上的MKLDNN和CUDNN中高度优化的基线相比,超过20倍的加速和90%的节能。因此,我们的方法开辟了一个新的研究方向,具有巨大的设备上培训潜力。11个代码:https://github.com/SLDGroup/GradientFilter-CVPR23
Despite its importance for federated learning, continuous learning and many other applications, on-device training remains an open problem for EdgeAI. The problem stems from the large number of operations (e.g., floating point multiplications and additions) and memory consumption required during training by the back-propagation algorithm. Consequently, in this paper, we propose a new gradient filtering approach which enables on-device CNN model training. More precisely, our approach creates a special structure with fewer unique elements in the gradient map, thus significantly reducing the computational complexity and memory consumption of back propagation during training. Extensive experiments on image classification and semantic segmentation with multiple CNN models (e.g., MobileNet, DeepLabV3, UPerNet) and devices (e.g., Raspberry Pi and Jetson Nano) demonstrate the effectiveness and wide applicability of our approach. For example, compared to SOTA, we achieve up to 19× speedup and 77.1% memory savings on ImageNet classification with only 0.1% accuracy loss. Finally, our method is easy to implement and deploy; over 20× speedup and 90% energy savings have been observed compared to highly optimized baselines in MKLDNN and CUDNN on NVIDIA Jetson Nano. Consequently, our approach opens up a new direction of research with a huge potential for on-device training.11Code: https://github.com/SLDGroup/GradientFilter-CVPR23