Heirarchical Federated Learning in Delay Sensitive Communication Networks

Heirarchical Federated Learning in Delay Sensitive Communication Networks
复制标题

DOI:
10.1109/ieeeconf56349.2022.10052096
复制
发表时间:
2022-10
期刊:
2022 56th Asilomar Conference on Signals, Systems, and Computers
影响因子:
--
通讯作者:
Abdulmoneam Ali;A. Arafa
Abdulmoneam Ali;A. Arafa
中科院分区:
其他
文献类型:
--
作者:
Abdulmoneam Ali;A. Arafa

文献摘要

相似文献

在客户端和参数服务器之间存在通信延迟的情况下,研究了局部平均对联邦学习系统性能的影响。我们关注分层FL(HFL)设置,其中客户端被分配到不同的组,每个组都有其自己的本地参数服务器(LPS)。我们工作中的本地和全局通信轮数由每组客户经历的(不同)延迟随机确定。具体地说,本地平均轮次的数量被绑定到被称为同步时间S的墙上时钟时间段,之后LPS通过与全球参数服务器(GPS)共享它们来同步它们的模型。然后重新应用这样的同步时间$S$,直到全球挂钟时间用完。首先,推导出在每个LPS处更新的模型相对于GPS处可用模型之间的偏差的上界。然后,这被用来推导出我们建议的HFL设置在每个LPS和在GPS处的收敛界限。结果表明,对于每组客户数量和延迟统计不同的异类系统,优化$S$尤为重要。除了为表现不佳的群体显示必要的协作需求外,优化$S$的值还促进了群体之间的公平,并允许人们处理对延迟敏感的FL申请,其中培训时间受到限制。
The impact of local averaging on the performance of federated learning (FL) systems is studied in the presence of communication delay between the clients and the parameter server. We focus on a hierarchical FL (HFL) setting where clients are assigned into different groups, each having its own local parameter server (LPS). The number of local and global communication rounds in our work is randomly determined by the (different) delays experienced by each group of clients. Specifically, the number of local averaging rounds are tied to a wall-clock time period coined the sync time S, after which the LPSs synchronize their models by sharing them with a global parameter server (GPS). Such sync time $S$ is then reapplied until a global wall-clock time is exhausted. First, an upper bound on the deviation between the updated model at each LPS with respect to that available at the GPS is derived. This is then used to derive the convergence bounds of our proposed HFL setting, at each LPS and at the GPS. The bounds showcase the effects of the whole system's parameters, including the number of groups, the number of clients per group, and the value of S. Our results show that optimizing $S$ is especially crucial in heterogeneous systems in which the number of clients per group and their delay statistics are different. In addition to showing the necessary need of collaboration for under-performing groups, optimizing the value of $S$ promotes fairness among groups, and allows one to deal with delay-sensitive FL applications in which the training time is restricted.