Fed-CBS: A Heterogeneity-Aware Client Sampling Mechanism for Federated Learning via Class-Imbalance Reduction

Fed-CBS: A Heterogeneity-Aware Client Sampling Mechanism for Federated Learning via Class-Imbalance Reduction
复制标题

DOI:
10.48550/arxiv.2209.15245
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Jianyi Zhang;Ang Li;Minxue Tang;Jingwei Sun;Xiang Chen;Fan Zhang;Chang Chen;Yiran Chen;H. Li
Jianyi Zhang;Ang Li;Minxue Tang;Jingwei Sun;Xiang Chen;Fan Zhang;Chang Chen;Yiran Chen;H. Li
中科院分区:
其他
文献类型:
--
作者:
Jianyi Zhang;Ang Li;Minxue Tang;Jingwei Sun;Xiang Chen;Fan Zhang;Chang Chen;Yiran Chen;H. Li

文献摘要

被引文献

相似文献

由于边缘设备的通信能力有限,大多数现有的联邦学习(FL)方法仅随机选择设备的子集来参与每个通信回合的训练。与使用所有可用客户端相比,随机选择机制可能会导致非IID(独立同分布)数据的性能显着下降。在本文中,我们展示了我们的关键观察,导致这种性能下降的根本原因是随机选择的客户端的分组数据的类不平衡。基于我们的关键观察,我们设计了一个有效的异构感知客户端采样机制,即,Federated Class-balanced Sampling(Fed-CBS),它可以有效地减少来自有意选择的客户端的组数据集的类不平衡。特别是,我们提出了一个措施的类不平衡,然后采用同态加密来获得这种措施的隐私保护的方式。在此基础上,我们还设计了一个计算效率高的客户端抽样策略,使积极选择的客户端将产生一个更类平衡的分组数据集与理论保证。大量的实验结果表明,Fed-CBS优于现状的方法。此外,它实现了与理想环境相当甚至更好的性能,在理想环境中,所有可用的客户都参与FL培训。
Due to limited communication capacities of edge devices, most existing federated learning (FL) methods randomly select only a subset of devices to participate in training for each communication round. Compared with engaging all the available clients, the random-selection mechanism can lead to significant performance degradation on non-IID (independent and identically distributed) data. In this paper, we show our key observation that the essential reason resulting in such performance degradation is the class-imbalance of the grouped data from randomly selected clients. Based on our key observation, we design an efficient heterogeneity-aware client sampling mechanism, i.e., Federated Class-balanced Sampling (Fed-CBS), which can effectively reduce class-imbalance of the group dataset from the intentionally selected clients. In particular, we propose a measure of class-imbalance and then employ homomorphic encryption to derive this measure in a privacy-preserving way. Based on this measure, we also design a computation-efficient client sampling strategy, such that the actively selected clients will generate a more class-balanced grouped dataset with theoretical guarantees. Extensive experimental results demonstrate Fed-CBS outperforms the status quo approaches. Furthermore, it achieves comparable or even better performance than the ideal setting where all the available clients participate in the FL training.