Straggler-Resilient Federated Learning: Leveraging the Interplay Between Statistical Accuracy and System Heterogeneity

Straggler-Resilient Federated Learning: Leveraging the Interplay Between Statistical Accuracy and System Heterogeneity
复制标题

DOI:
10.1109/jsait.2022.3205475
复制
发表时间:
2020-12
期刊:
IEEE Journal on Selected Areas in Information Theory
影响因子:
--
通讯作者:
Amirhossein Reisizadeh;Isidoros Tziotis;Hamed Hassani;Aryan Mokhtari;Ramtin Pedarsani
Amirhossein Reisizadeh;Isidoros Tziotis;Hamed Hassani;Aryan Mokhtari;Ramtin Pedarsani
中科院分区:
其他
文献类型:
--
作者:
Amirhossein Reisizadeh;Isidoros Tziotis;Hamed Hassani;Aryan Mokhtari;Ramtin Pedarsani

文献摘要

被引文献

相似文献

联邦学习是一种新颖的范式,涉及从数据样本中分发到大型客户网络的数据,而数据仍然是本地的。能力。由于设备缓慢(Stragglers)的运行时间,我们提出了一种新颖的散乱的联盟学习元学习元元素,以适应客户的统计特征,以便适应客户的统计特征提高学习过程。到达节点的数据,而每个阶段的最终模型被用作下一阶段的温暖启动模型。 I.I.D.针对特定实例,Flanp将整体预期运行时削减$ \ Mathcal {O}(\ ln(ns))$,其中$ n $和$ s $表示。与标准的联邦学习基准相比,在实验中,每个节点的样本分别显示出明显的加速时间 - $ 6 \ times $。
Federated learning is a novel paradigm that involves learning from data samples distributed across a large network of clients while the data remains local. It is, however, known that federated learning is prone to multiple system challenges including system heterogeneity where clients have different computation and communication capabilities. Such heterogeneity in clients’ computation speed has a negative effect on the scalability of federated learning algorithms and causes significant slow-down in their runtime due to slow devices (stragglers). In this paper, we propose FLANP, a novel straggler-resilient federated learning meta-algorithm that incorporates statistical characteristics of the clients’ data to adaptively select the clients in order to speed up the learning procedure. The key idea of FLANP is to start the training procedure with faster nodes and gradually involve the slower ones in the model training once the statistical accuracy of the current participating nodes’ data is reached, while the final model for each stage is used as a warm-start model for the next stage. Our theoretical results characterize the speedup provided by the meta-algorithm FLANP in comparison to standard federated benchmarks for strongly convex losses and i.i.d. samples. For particular instances, FLANP slashes the overall expected runtime by a factor of $\mathcal {O}(\ln (Ns))$ , where $N$ and $s$ denote the total number of nodes and the number of samples per node, respectively. In experiments, FLANP demonstrates significant speedups in wall-clock time -up to $6 \times $ – compared to standard federated learning benchmarks.