Federated Learning with Non-IID Data

Federated Learning with Non-IID Data
复制标题

DOI:
10.48550/arxiv.1806.00582
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yue Zhao;Meng Li;Liangzhen Lai;Naveen Suda;Damon Civin;V. Chandra
Yue Zhao;Meng Li;Liangzhen Lai;Naveen Suda;Damon Civin;V. Chandra
中科院分区:
其他
文献类型:
--
作者:
Yue Zhao;Meng Li;Liangzhen Lai;Naveen Suda;Damon Civin;V. Chandra

文献摘要

被引文献

相似文献

联合学习使资源受限的边缘计算设备(如移动的电话和IoT设备)能够学习用于预测的共享模型,同时将训练数据保持在本地。这种去中心化的训练模型方法提供了隐私、安全、监管和经济效益。在这项工作中,我们专注于当本地数据是非IID时联邦学习的统计挑战。我们首先表明,联邦学习的准确性显着降低,对于针对高度偏斜的非IID数据训练的神经网络,其准确性最高可降低55%,其中每个客户端设备仅对单一类别的数据进行训练。我们进一步表明,这种精度降低可以解释的重量分歧,这可以量化的推土机的距离(EMD)之间的分布在每个设备上的类和人口分布。作为一种解决方案,我们提出了一种策略,通过创建一个在所有边缘设备之间全局共享的小数据子集来改进非IID数据的训练。实验表明,CIFAR-10数据集的准确率可以提高30%,只有5%的全球共享数据。
Federated learning enables resource-constrained edge compute devices, such as mobile phones and IoT devices, to learn a shared model for prediction, while keeping the training data local. This decentralized approach to train models provides privacy, security, regulatory and economic benefits. In this work, we focus on the statistical challenge of federated learning when local data is non-IID. We first show that the accuracy of federated learning reduces significantly, by up to 55% for neural networks trained for highly skewed non-IID data, where each client device trains only on a single class of data. We further show that this accuracy reduction can be explained by the weight divergence, which can be quantified by the earth mover's distance (EMD) between the distribution over classes on each device and the population distribution. As a solution, we propose a strategy to improve training on non-IID data by creating a small subset of data which is globally shared between all the edge devices. Experiments show that accuracy can be increased by 30% for the CIFAR-10 dataset with only 5% globally shared data.