IoT Device Friendly and Communication-Efficient Federated Learning via Joint Model Pruning and Quantization

IoT Device Friendly and Communication-Efficient Federated Learning via Joint Model Pruning and Quantization
复制标题

DOI:
10.1109/jiot.2022.3145865
复制
发表时间:
2022-08-01
影响因子:
10.6
通讯作者:
Pan, Miao
Pan, Miao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Prakash, Pavana;Ding, Jiahao;Pan, Miao

文献摘要

被引文献

相似文献

联邦学习(FL)通过其新颖的应用和服务增强了其作为物联网(IoT)领域有前途的工具的存在。具体来说,在具有大量物联网设备的多访问边缘计算设置中,FL是最合适的,因为它利用分布式客户端数据来训练高性能深度学习(DL)模型,同时保持数据的私密性。然而,底层的深度神经网络(dnn)是巨大的,阻止了它直接部署到资源有限的计算和内存有限的物联网设备上。此外,在FL中,中央服务器和客户端之间频繁交换模型更新可能导致通信瓶颈。为了解决这些挑战,在本文中,我们介绍了GWEP,一种基于模型压缩的FL方法。它利用联合量化和模型修剪来获得深度神经网络的优势,同时满足资源受限设备的能力。因此,通过减少FL的计算、内存和网络占用,低端物联网设备可能能够参与FL过程。此外,我们还提供了FL收敛性的理论保证。通过实证评估,我们证明我们的方法显著优于基线算法,速度高达10.23倍,通信回合减少11倍,同时实现高模型压缩、能源效率和学习性能。
Federated learning (FL) through its novel applications and services has enhanced its presence as a promising tool in the Internet of Things (IoT) domain. Specifically, in a multiaccess edge computing setup with a host of IoT devices, FL is most suitable since it leverages distributed client data to train high-performance deep learning (DL) models while keeping the data private. However, the underlying deep neural networks (DNNs) are huge, preventing its direct deployment onto resource-constrained computing and memory-limited IoT devices. Besides, frequent exchange of model updates between the central server and clients in FL could result in a communication bottleneck. To address these challenges, in this article, we introduce GWEP, a model compression-based FL method. It utilizes joint quantization and model pruning to reap the benefits of DNNs while meeting the capabilities of resource-constrained devices. Consequently, by reducing the computational, memory, and network footprint of FL, the low-end IoT devices may be able to participate in the FL process. In addition, we provide theoretical guarantees of FL convergence. Through empirical evaluations, we demonstrate that our approach significantly outperforms the baseline algorithms by being up to 10.23 times faster with 11 times lesser communication rounds, while achieving high-model compression, energy efficiency, and learning performance.