FedLGA: Toward System-Heterogeneity of Federated Learning via Local Gradient Approximation

FedLGA: Toward System-Heterogeneity of Federated Learning via Local Gradient Approximation
复制标题

DOI:
10.1109/tcyb.2023.3247365
复制
发表时间:
2021-12
影响因子:
11.8
通讯作者:
Xingyu Li;Zhe Qu;Bo Tang;Zhuo Lu
Xingyu Li;Zhe Qu;Bo Tang;Zhuo Lu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Xingyu Li;Zhe Qu;Bo Tang;Zhuo Lu

文献摘要

相似文献

联邦学习(FL)是一种去中心化的机器学习架构,它利用大量远程设备来学习具有分布式训练数据的联合模型。然而,系统异构性是 FL 网络实现鲁棒分布式学习性能的一大挑战,它来自两个方面:1)由于设备之间的计算能力不同而导致的设备异构性;2)由于网络上分布的数据不相同而导致的数据异构性。先前针对异构 FL 问题的研究(例如 FedProx)缺乏形式化,并且仍然是一个悬而未决的问题。这项工作首先形式化了系统异构 FL 问题,并提出了一种称为联合局部梯度近似 (FedLGA) 的新算法,通过梯度近似弥合局部模型更新的分歧来解决该问题。为了实现这一点,FedLGA 提供了一种替代 Hessian 估计方法,该方法只需要聚合器上有额外的线性复杂度。理论上,我们表明,通过设备异构比率 $\rho $ ,FedLGA 可以实现非独立同分布的收敛率。非凸优化问题的分布式 FL 训练数据,分别用于完全和部分设备参与,其中 $\mathcal{O} ({}[{(1+\rho)}/{\sqrt {ENT}}] + {}{1}/{T})$ 和 $\mathcal{O} ({}[{(1+\rho)\sqrt {E}}/{\sqrt {TK}}] + {}{1}/{T})$,其中$E$是本地学习epoch数,$T$是总通信轮数,$N$是总设备数,$K$是部分参与方案下一轮通信中选定的设备数。在多个数据集上的综合实验结果表明,FedLGA 可以有效解决系统异构问题,并且优于当前的 FL 方法。具体来说,针对 CIFAR-10 数据集的性能表明,与 FedAvg 相比,FedLGA 将模型的最佳测试精度从 60.91% 提高到 64.44%。
Federated learning (FL) is a decentralized machine learning architecture, which leverages a large number of remote devices to learn a joint model with distributed training data. However, the system-heterogeneity is one major challenge in an FL network to achieve robust distributed learning performance, which comes from two aspects: 1) device-heterogeneity due to the diverse computational capacity among devices and 2) data-heterogeneity due to the nonidentically distributed data across the network. Prior studies addressing the heterogeneous FL issue, for example, FedProx, lack formalization and it remains an open problem. This work first formalizes the system-heterogeneous FL problem and proposes a new algorithm, called federated local gradient approximation (FedLGA), to address this problem by bridging the divergence of local model updates via gradient approximation. To achieve this, FedLGA provides an alternated Hessian estimation method, which only requires extra linear complexity on the aggregator. Theoretically, we show that with a device-heterogeneous ratio $\rho $ , FedLGA achieves convergence rates on non-i.i.d. distributed FL training data for the nonconvex optimization problems with $\mathcal{O} ({}[{(1+\rho)}/{\sqrt {ENT}}] + {}{1}/{T})$ and $\mathcal{O} ({}[{(1+\rho)\sqrt {E}}/{\sqrt {TK}}] + {}{1}/{T})$ for full and partial device participation, respectively, where $E$ is the number of local learning epoch, $T$ is the number of total communication round, $N$ is the total device number, and $K$ is the number of the selected device in one communication round under partially participation scheme. The results of comprehensive experiments on multiple datasets indicate that FedLGA can effectively address the system-heterogeneous problem and outperform current FL methods. Specifically, the performance against the CIFAR-10 dataset shows that, compared with FedAvg, FedLGA improves the model’s best testing accuracy from 60.91% to 64.44%.