Sparse Communication for Federated Learning

Sparse Communication for Federated Learning
复制标题

DOI:
10.1109/icfec54809.2022.00008
复制
发表时间:
2022-05
期刊:
2022 IEEE 6th International Conference on Fog and Edge Computing (ICFEC)
影响因子:
--
通讯作者:
Kundjanasith Thonglek;Keichi Takahashi;Koheix Ichikawa;Chawanat Nakasan;P. Leelaprute;Hajimu Iida
Kundjanasith Thonglek;Keichi Takahashi;Koheix Ichikawa;Chawanat Nakasan;P. Leelaprute;Hajimu Iida
中科院分区:
其他
文献类型:
--
作者:
Kundjanasith Thonglek;Keichi Takahashi;Koheix Ichikawa;Chawanat Nakasan;P. Leelaprute;Hajimu Iida

文献摘要

相似文献

联合学习使用在大量边缘设备上分布的数据集在集中式服务器上训练模型。由于Federated Learning不会将本地数据从边缘设备发送到服务器,因此可以保留数据隐私。它从边缘设备而不是本地数据传输本地模型。但是,沟通成本通常是联邦学习中的问题。本文提出了一种新的方法,以通过仅传输神经网络模型中的顶级更新参数来降低联合学习所需的沟通成本。提出的方法允许调整更新参数的标准,以权衡降低通信成本和模型准确性的损失。我们使用各种模型和数据集评估了提出的方法,发现它可以实现可比的性能,以转移用于联合学习的原始模型。结果,与传统的VGG16方法相比,所提出的方法已降低了所需的通信成本约为90%。此外,我们发现所提出的方法能够降低大型模型的通信成本,而不是小型模型,因为每个模型体系结构中更新参数的阈值不同。
Federated learning trains a model on a centralized server using datasets distributed over a massive amount of edge devices. Since federated learning does not send local data from edge devices to the server, it preserves data privacy. It transfers the local models from edge devices instead of the local data. However, communication costs are frequently a problem in federated learning. This paper proposes a novel method to reduce the required communication cost for federated learning by transferring only top updated parameters in neural network models. The proposed method allows adjusting the criteria of updated parameters to trade-off the reduction of communication costs and the loss of model accuracy. We evaluated the proposed method using diverse models and datasets and found that it can achieve comparable performance to transfer original models for federated learning. As a result, the proposed method has achieved a reduction of the required communication costs around 90% when compared to the conventional method for VGG16. Furthermore, we found out that the proposed method is able to reduce the communication cost of a large model more than of a small model due to the different threshold of updated parameters in each model architecture.