Convergence of Update Aware Device Scheduling for Federated Learning at the Wireless Edge

Convergence of Update Aware Device Scheduling for Federated Learning at the Wireless Edge
复制标题

DOI:
10.1109/twc.2021.3052681
复制
发表时间:
2021-06-01
影响因子:
10.4
通讯作者:
Poor, H. Vincent
Poor, H. Vincent
中科院分区:
计算机科学1区
文献类型:
--
作者:
Amiri, Mohammad Mohammadi;Gunduz, Deniz;Poor, H. Vincent

文献摘要

被引文献

相似文献

我们研究了无线边缘的联合学习(FL),在这种学习中,具有本地数据集的功率受限设备在远程参数服务器(PS)的帮助下协作训练联合模型。我们假设这些设备通过带宽有限的共享无线信道连接到PS。在FL的每次迭代中,设备的子集被调度以在正交信道资源上将其本地模型更新发送到PS,而每个参与设备必须压缩其模型更新以适应其链路容量。我们设计了新颖的调度和资源分配策略,不仅基于参与设备的信道条件,而且基于其本地模型更新的重要性,来决定在每一轮中要传输的设备的子集,以及如何在参与设备之间分配资源。然后,我们建立了无线FL算法与设备调度的收敛,其中设备传送其消息的能力有限。数值实验结果表明,基于信道条件和局部模型更新重要性的调度策略比单独基于这两个度量之一的调度策略具有更好的长期性能。此外,我们观察到,当数据是独立且同分布的(I.I.D.)在设备之间,在每轮中选择一个设备可提供最佳性能,而当数据分布为非I.I.D.时,在每轮中调度多个设备可提高性能。这一观察结果得到了融合结果的验证,该结果表明,对于多样性较小、偏向较大的数据分布,调度设备的数量应该增加。
We study federated learning (FL) at the wireless edge, where power-limited devices with local datasets collaboratively train a joint model with the help of a remote parameter server (PS). We assume that the devices are connected to the PS through a bandwidth-limited shared wireless channel. At each iteration of FL, a subset of the devices are scheduled to transmit their local model updates to the PS over orthogonal channel resources, while each participating device must compress its model update to accommodate to its link capacity. We design novel scheduling and resource allocation policies that decide on the subset of the devices to transmit at each round, and how the resources should be allocated among the participating devices, not only based on their channel conditions, but also on the significance of their local model updates. We then establish convergence of a wireless FL algorithm with device scheduling, where devices have limited capacity to convey their messages. The results of numerical experiments show that the proposed scheduling policy, based on both the channel conditions and the significance of the local model updates, provides a better long-term performance than scheduling policies based only on either of the two metrics individually. Furthermore, we observe that when the data is independent and identically distributed (i.i.d.) across devices, selecting a single device at each round provides the best performance, while when the data distribution is non-i.i.d., scheduling multiple devices at each round improves the performance. This observation is verified by the convergence result, which shows that the number of scheduled devices should increase for a less diverse and more biased data distribution.