Anarchic Federated learning with Delayed Gradient Averaging

Anarchic Federated learning with Delayed Gradient Averaging
复制标题

DOI:
10.1145/3565287.3610273
复制
发表时间:
2023-10
期刊:
Proceedings of the Twenty-fourth International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing
影响因子:
--
通讯作者:
Dongsheng Li;Xiaowen Gong
Dongsheng Li;Xiaowen Gong
中科院分区:
其他
文献类型:
--
作者:
Dongsheng Li;Xiaowen Gong

文献摘要

相似文献

联邦学习(FL)在过去几年中的快速发展最近激发了对这一新兴主题的大量研究。FL的现有工作通常假设客户以某种特定的模式(如平衡参与)和/或以同步方式和/或以相同数量的局部迭代参与学习过程,而这些假设在实践中可能很难保持。在本文中,我们提出了AFL-DGA,一种具有延迟梯度平均的无政府联邦学习算法,它为客户端提供了最大的自由度。特别是,AFL-DGA允许客户端1)参与任何轮次; 2)异步参与; 3)参与任何数量的局部迭代; 4)并行执行梯度计算和梯度通信。AFL-DGA算法使客户端能够根据其异构和时变的计算和通信能力灵活地参与FL,并有效地提高其计算和通信资源的利用率。我们表征性能界限的学习损失的AFL-DGA作为客户端的本地迭代次数,本地模型延迟,和全球模型延迟的函数。我们的研究结果表明,AFL-DGA算法可以实现[方程]的收敛速度和线性收敛加速,这与现有的基准相匹配。研究结果还描述了各种系统参数对学习损失的影响,这提供了有用的见解。数值结果表明了该算法的有效性。
The rapid advances in federated learning (FL) in the past few years have recently inspired a great deal of research on this emerging topic. Existing work on FL often assume that clients participate in the learning process with some particular pattern (such as balanced participation), and/or in a synchronous manner, and/or with the same number of local iterations, while these assumptions can be hard to hold in practice. In this paper, we propose AFL-DGA, an Anarchic Federated Learning algorithm with Delayed Gradient Averaging, which gives maximum freedom to clients. In particular, AFL-DGA allows clients to 1) participate in any rounds; 2) participate asynchronously; 3) participate with any number of local iterations; 4) perform gradient computations and gradient communications in parallel. The proposed AFL-DGA algorithm enables clients to participate in FL flexibly according to their heterogeneous and time-varying computation and communication capabilities, and also efficiently by improving utilization of their computation and communication resources. We characterize performance bounds on the learning loss of AFL-DGA as a function of clients' local iteration numbers, local model delays, and global model delays. Our results show that the AFL-DGA algorithm can achieve a convergence rate of [EQUATION] and also a linear convergence speedup, which matches that of existing benchmarks. The results also characterize the impacts of various system parameters on the learning loss, which provide useful insights. Numerical results demonstrate the efficiency of the proposed algorithm.