AdaCoOpt: Leverage the Interplay of Batch Size and Aggregation Frequency for Federated Learning

AdaCoOpt: Leverage the Interplay of Batch Size and Aggregation Frequency for Federated Learning
复制标题

DOI:
10.1109/iwqos57198.2023.10188807
复制
发表时间:
2023-06
期刊:
2023 IEEE/ACM 31st International Symposium on Quality of Service (IWQoS)
影响因子:
--
通讯作者:
Weijie Liu-;Xiaoxi Zhang;Jingpu Duan;Carlee Joe-Wong;Zhi Zhou;Xu Chen
Weijie Liu-;Xiaoxi Zhang;Jingpu Duan;Carlee Joe-Wong;Zhi Zhou;Xu Chen
中科院分区:
其他
文献类型:
--
作者:
Weijie Liu-;Xiaoxi Zhang;Jingpu Duan;Carlee Joe-Wong;Zhi Zhou;Xu Chen

文献摘要

相似文献

联合学习(FL)是一种分布式学习范式,可以协调异构边缘设备来执行模型训练,而无需共享私有原始数据。许多先前的工作已经分析了FL收敛相对于重要的超参数,包括批量大小和聚合频率。然而,调整批量大小和本地更新的数量可能会以不同且可能复杂的形式影响模型性能、训练时间以及消耗计算和通信资源的成本。它们的联合效应被忽视了,应该加以利用,以实现精确的模型与可控的业务支出。本文提出了新的分析模型和优化算法,利用批量大小和聚合频率的相互作用来导航FL的收敛性,成本和完成时间之间的权衡。我们首先在跨设备的异构训练数据集下获得了一个新的训练误差收敛界。基于这个界限,我们得到了封闭形式的解决方案的共同优化的批量大小和聚合频率,一个单一的配置为所有的设备。然后,我们设计了一个高效的精确算法,用于跨设备分配不同的批处理配置,可以进一步提高模型的准确性,以解决数据和系统特性的异构性。此外,我们提出了一个自适应控制算法,动态调整的解决方案与估计的网络状态。大量的实验证明了我们的离线最优解和在线自适应算法的优越性。
Federated Learning (FL) is a distributed learning paradigm that can coordinate heterogeneous edge devices to perform model training without sharing private raw data. Many prior works have analyzed the FL convergence with respect to important hyperparameters, including batch size and aggregation frequency. However, adjusting the batch size and the number of local updates can affect the model performance, training time, and the cost of consuming computation and communication resources, in different and perhaps complex forms. Their joint effects have been overlooked and should be exploited to achieve accurate models with controllable operational expenditure. This paper proposes novel analytical models and optimization algorithms that leverage the interplay of batch size and aggregation frequency to navigate the trade-offs among convergence, cost, and completion time for FL. We first obtain a new convergence bound of the training error under heterogeneous training datasets across devices. Based on this bound, we derive closed-form solutions of a co-optimized batch size and aggregation frequency, a single configuration for all the devices. We then design an efficient exact algorithm for assigning different batch configurations across devices that can further improve the model accuracy to address the heterogeneity of both data and system characteristics. Further, we propose an adaptive control algorithm to dynamically adjust the solutions with estimated network states. Extensive experiments demonstrate the superiority of our offline optimal solutions and online adaptive algorithm.