Group Knowledge Transfer: Federated Learning of Large CNNs at the Edge

Group Knowledge Transfer: Federated Learning of Large CNNs at the Edge
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
arXiv: Learning
影响因子:
--
通讯作者:
Chaoyang He;M. Annavaram;S. Avestimehr
Chaoyang He;M. Annavaram;S. Avestimehr
中科院分区:
其他
文献类型:
--
作者:
Chaoyang He;M. Annavaram;S. Avestimehr

文献摘要

被引文献

相似文献

放大卷积神经网络(CNN)大小(例如,宽度、深度等)可以有效地提高模型精度。然而,大的模型尺寸阻碍了在资源受限的边缘设备上的训练。例如,联邦学习(FL)可能会给边缘节点的计算能力带来不必要的负担,尽管由于其隐私和机密性,对FL有很强的实际需求。为了解决边缘设备资源受限的现实,我们重新制定FL作为一个组的知识转移训练算法,称为FedGKT。FedGKT设计了一种交替最小化方法的变体,用于在边缘节点上训练小型CNN,并通过知识蒸馏将其知识定期转移到大型服务器端CNN。FedGKT将几个优势整合到一个框架中:减少对边缘计算的需求,降低大型CNN的通信带宽,以及异步训练,同时保持与FedAvg相当的模型准确性。我们使用三个不同的数据集(CIFAR-10,CIFAR-100和CINIC-10)及其非I.I. D训练基于ResNet-56和ResNet-110设计的CNN。变体。我们的研究结果表明,FedGKT可以获得可比的,甚至略高于FedAvg的准确性。更重要的是,FedGKT使边缘培训变得负担得起。与使用FedAvg的边缘训练相比,FedGKT对边缘设备的计算能力(FLOPs)要求降低9到17倍,对边缘CNN的参数要求降低54到105倍。我们的源代码发布在FedML(这个https URL)。
Scaling up the convolutional neural network (CNN) size (e.g., width, depth, etc.) is known to effectively improve model accuracy. However, the large model size impedes training on resource-constrained edge devices. For instance, federated learning (FL) may place undue burden on the compute capability of edge nodes, even though there is a strong practical need for FL due to its privacy and confidentiality properties. To address the resource-constrained reality of edge devices, we reformulate FL as a group knowledge transfer training algorithm, called FedGKT. FedGKT designs a variant of the alternating minimization approach to train small CNNs on edge nodes and periodically transfer their knowledge by knowledge distillation to a large server-side CNN. FedGKT consolidates several advantages into a single framework: reduced demand for edge computation, lower communication bandwidth for large CNNs, and asynchronous training, all while maintaining model accuracy comparable to FedAvg. We train CNNs designed based on ResNet-56 and ResNet-110 using three distinct datasets (CIFAR-10, CIFAR-100, and CINIC-10) and their non-I.I.D. variants. Our results show that FedGKT can obtain comparable or even slightly higher accuracy than FedAvg. More importantly, FedGKT makes edge training affordable. Compared to the edge training using FedAvg, FedGKT demands 9 to 17 times less computational power (FLOPs) on edge devices and requires 54 to 105 times fewer parameters in the edge CNN. Our source code is released at FedML (this https URL).