Robust Federated Learning for Unreliable and Resource-limited Wireless Networks

Robust Federated Learning for Unreliable and Resource-limited Wireless Networks
复制标题

DOI:
10.1109/twc.2024.3366393
复制
发表时间:
2024
影响因子:
10.4
通讯作者:
Zhixiong Chen;Wenqiang Yi;Yuanwei Liu;Arumgam Nallanathan
Zhixiong Chen;Wenqiang Yi;Yuanwei Liu;Arumgam Nallanathan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhixiong Chen;Wenqiang Yi;Yuanwei Liu;Arumgam Nallanathan

文献摘要

被引文献

相似文献

联邦学习(FL)是一种高效且保护隐私的分布式学习范式,使大规模边缘设备能够协同训练机器学习模型。尽管已经提出了各种通信方案来加速资源有限的无线网络中的FL过程,但对无线信道的不可靠性质的探索较少。在这项工作中,我们提出了一种新的FL框架,即FL与梯度回收(FL-GR),它消除了历史梯度的非调度和传输失败的设备,以提高FL的学习性能。为了减少硬件要求,实现FL-GR在实际网络中,我们开发了一个内存友好的FL-GR,相当于FL-GR,但需要低内存的边缘服务器。然后从理论上分析了无线网络参数对FL-GR收敛界的影响,发现最小化局部梯度失效的平均平方(AS-GS)有助于提高学习性能。在此基础上,我们制定了一个联合的设备调度,资源分配和功率控制优化问题,以最小化全局损耗最小化的AS-GS。为了解决这个问题,我们首先推导出设备的最优功率控制策略,并将AS-GS最小化问题转化为二分图匹配问题。通过详细的分析,我们进一步将二部匹配问题转化为一个等价的线性规划,便于求解。在三个真实世界数据集上的广泛模拟结果(即,MNIST、CIFAR-10和CIFAR-100)艾德了所提出方法的有效性。与不带梯度循环的FL算法相比,FL-GR算法具有更高的精度和更快的收敛速度。此外,本文提出的设备调度和资源分配算法在准确性和收敛速度方面也优于基准测试。
—Federated learning (FL) is an efficient and privacy-preserving distributed learning paradigm that enables massive edge devices to train machine learning models collaboratively. Although various communication schemes have been proposed to expedite the FL process in resource-limited wireless networks, the unreliable nature of wireless channels was less explored. In this work, we propose a novel FL framework, namely FL with gradient recycling (FL-GR), which recycles the historical gradients of unscheduled and transmission-failure devices to improve the learning performance of FL. To reduce the hardware requirements for implementing FL-GR in the practical network, we develop a memory-friendly FL-GR that is equivalent to FL-GR but requires low memory of the edge server. We then theoretically analyze how the wireless network parameters affect the convergence bound of FL-GR, revealing that minimizing the average square of local gradients’ staleness (AS-GS) helps improve the learning performance. Based on this, we formulate a joint device scheduling, resource allocation and power control optimization problem to minimize the AS-GS for global loss minimization. To solve the problem, we first derive the optimal power control policy for devices and transform the AS-GS minimization problem into a bipartite graph matching problem. Through detailed analysis, we further transform the bipartite matching problem into an equivalent linear program which is convenient to solve. Extensive simulation results on three real-world datasets (i.e., MNIST, CIFAR-10, and CIFAR-100) verified the efficacy of the proposed methods. Compared to the FL algorithms without gradient recycling, FL-GR is able to achieve higher accuracy and fast convergence speed. In addition, the proposed device scheduling and resource allocation algorithm also outperforms the benchmarks in accuracy and convergence speed.