HACCS: Heterogeneity-Aware Clustered Client Selection for Accelerated Federated Learning

HACCS: Heterogeneity-Aware Clustered Client Selection for Accelerated Federated Learning
复制标题

DOI:
10.1109/ipdps53621.2022.00100
复制
发表时间:
2022-05
期刊:
2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Joel Wolfrath;N. Sreekumar;Dhruv Kumar;Yuanli Wang;A. Chandra
Joel Wolfrath;N. Sreekumar;Dhruv Kumar;Yuanli Wang;A. Chandra
中科院分区:
其他
文献类型:
--
作者:
Joel Wolfrath;N. Sreekumar;Dhruv Kumar;Yuanli Wang;A. Chandra

文献摘要

相似文献

联合学习是一种机器学习范式,其中全局模型在大量分布式边缘设备上进行原位训练。虽然这种技术避免了将数据传输到中央位置的成本,并实现了高度的隐私,但由于可用于训练的异构硬件资源,它带来了额外的挑战。此外,数据在所有边缘设备之间并不是独立和相同分布的(IID),导致设备之间的统计异质性。由于这些限制,客户端选择策略在模型训练期间对及时收敛起着重要作用。现有的策略确保每个单独的设备至少定期地包括在培训过程中。在这项工作中,我们提出了HACCS,一个异质性感知的客户端选择系统,识别和利用的统计异质性,通过代表所有可区分的数据分布,而不是在训练过程中的个别设备。如果系统中的其他设备具有类似的数据分布,则HACCS对单个设备的丢失具有鲁棒性。我们提出了隐私保护的方法来估计这些客户端分布和聚类。我们还提出了利用这些集群在联邦学习系统中做出调度决策的策略。我们对真实世界数据集的评估表明,与最先进的技术相比,我们的框架可以在不影响准确性的情况下减少18%-38%的收敛时间。
Federated Learning is a machine learning paradigm where a global model is trained in-situ across a large number of distributed edge devices. While this technique avoids the cost of transferring data to a central location and achieves a strong degree of privacy, it presents additional challenges due to the heterogeneous hardware resources available for training. Furthermore, data is not independent and identically distributed (IID) across all edge devices, resulting in statistical heterogeneity across devices. Due to these constraints, client selection strategies play an important role for timely convergence during model training. Existing strategies ensure that each individual device is included, at least periodically, in the training process. In this work, we propose HACCS, a Heterogeneity-Aware Clustered Client Selection system that identifies and exploits the statistical heterogeneity by representing all distinguishable data distributions instead of individual devices in the training process. HACCS is robust to individual device dropout, provided other devices in the system have similar data distributions. We propose privacy-preserving methods for estimating these client distributions and clustering them. We also propose strategies for leveraging these clusters to make scheduling decisions in a federated learning system. Our evaluation on real-world datasets suggests that our framework can provide 18% −38% reduction in time to convergence compared to the state of the art without any compromise in accuracy.