FedAT: A High-Performance and Communication-Efficient Federated Learning System with Asynchronous Tiers

FedAT: A High-Performance and Communication-Efficient Federated Learning System with Asynchronous Tiers
复制标题

DOI:
10.1145/3458817.3476211
复制
发表时间:
2020-10
期刊:
SC21: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Zheng Chai;Yujing Chen;Ali Anwar;Liang Zhao;Yue Cheng;H. Rangwala
Zheng Chai;Yujing Chen;Ali Anwar;Liang Zhao;Yue Cheng;H. Rangwala
中科院分区:
其他
文献类型:
--
作者:
Zheng Chai;Yujing Chen;Ali Anwar;Liang Zhao;Yue Cheng;H. Rangwala

文献摘要

被引文献

相似文献

联邦学习(FL)涉及在大规模分布式设备上训练模型,同时保持训练数据的本地化和私有化。这种形式的协作学习暴露了模型收敛速度,模型精度,跨客户端的平衡和通信成本之间的新的权衡,新的挑战包括:(1)掉队的问题,其中客户端由于数据或(计算和网络)资源的异构性而滞后,以及(2)通信拥塞,其中大量的客户端将其本地更新通信到中央服务器并使服务器瓶颈。许多现有的FL方法专注于优化沿着只有一个单一的维度的权衡空间。现有的解决方案使用异步模型更新或基于分层的同步机制来解决落伍者问题。然而,异步方法很容易造成通信瓶颈,而分层可能会引入偏向,倾向于更快的层和更短的响应时间。为了解决这些问题,我们提出了FedAT,一个新的联邦学习系统与异步层下非i.i.d.训练数据FedAT协同地结合了同步的层内培训和异步的跨层培训。通过分层桥接同步和异步训练,FedAT最大限度地减少了落伍者的影响,提高了收敛速度和测试精度。FedAT使用一种能够感知掉队者的加权聚合启发式方法来引导和平衡客户端之间的训练,以进一步提高准确性。FedAT使用高效的基于折线编码的压缩算法压缩上行链路和下行链路通信,从而最大限度地降低通信成本。结果表明,FedAT的预测性能提高了21.09%,并减少了高达8.5倍的通信成本,相比最先进的FL方法。
Federated learning (FL) involves training a model over massive distributed devices, while keeping the training data localized and private. This form of collaborative learning exposes new tradeoffs among model convergence speed, model accuracy, balance across clients, and communication cost, with new challenges including: (1) straggler problem-where clients lag due to data or (computing and network) resource heterogeneity, and (2) communication bottleneck-where a large number of clients communicate their local updates to a central server and bottleneck the server. Many existing FL methods focus on optimizing along only one single dimension of the tradeoff space. Existing solutions use asynchronous model updating or tiering-based, synchronous mechanisms to tackle the straggler problem. However, asynchronous methods can easily create a communication bottleneck, while tiering may introduce biases that favor faster tiers with shorter response latencies. To address these issues, we present FedAT, a novel Federated learning system with Asynchronous Tiers under Non-i.i.d. training data. FedAT synergistically combines synchronous, intra-tier training and asynchronous, cross-tier training. By bridging the synchronous and asynchronous training through tiering, FedAT minimizes the straggler effect with improved convergence speed and test accuracy. FedAT uses a straggler-aware, weighted aggregation heuristic to steer and balance the training across clients for further accuracy improvement. FedAT compresses uplink and downlink communications using an efficient, polyline-encoding-based compression algorithm, which minimizes the communication cost. Results show that FedAT improves the prediction performance by up to 21.09% and reduces the communication cost by up to 8.5×, compared to state-of-the-art FL methods.