QuPeD: Quantized Personalization via Distillation with Applications to Federated Learning

QuPeD: Quantized Personalization via Distillation with Applications to Federated Learning
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
Kaan Ozkara;Navjot Singh;Deepesh Data;S. Diggavi
Kaan Ozkara;Navjot Singh;Deepesh Data;S. Diggavi
中科院分区:
其他
文献类型:
--
作者:
Kaan Ozkara;Navjot Singh;Deepesh Data;S. Diggavi

文献摘要

相似文献

传统上,联邦学习 (FL) 的目标是在协作使用多个客户端和服务器的同时训练单个全局模型。 FL 算法面临的两个自然挑战是跨客户端数据的异构性以及客户端与{\em不同资源}的协作。在这项工作中,我们引入了 \textit{量化} 和 \textit{个性化} FL 算法 QuPeD,该算法通过 \textit{知识蒸馏} (KD) 在有权访问异构数据和资源的客户端之间促进集体(个性化模型压缩)训练。对于个性化,我们允许客户学习具有不同量化参数和模型维度/结构的\textit{压缩的个性化模型}。为此,首先我们提出了一种通过宽松的优化问题来学习量化模型的算法,其中量化值也得到了优化。当参与(联邦)学习过程的每个客户对压缩模型有不同的要求(模型维度和精度)时,我们通过引入通过全局模型协作的本地客户目标的知识蒸馏损失来制定压缩个性化框架。我们开发了一种交替近端梯度更新来解决这个压缩个性化问题,并分析其收敛特性。在数值上,我们验证了 QuPeD 优于竞争性个性化 FL 方法、FedAvg 以及在各种异构环境中对客户进行的本地培训。
Traditionally, federated learning (FL) aims to train a single global model while collaboratively using multiple clients and a server. Two natural challenges that FL algorithms face are heterogeneity in data across clients and collaboration of clients with {\em diverse resources}. In this work, we introduce a \textit{quantized} and \textit{personalized} FL algorithm QuPeD that facilitates collective (personalized model compression) training via \textit{knowledge distillation} (KD) among clients who have access to heterogeneous data and resources. For personalization, we allow clients to learn \textit{compressed personalized models} with different quantization parameters and model dimensions/structures. Towards this, first we propose an algorithm for learning quantized models through a relaxed optimization problem, where quantization values are also optimized over. When each client participating in the (federated) learning process has different requirements for the compressed model (both in model dimension and precision), we formulate a compressed personalization framework by introducing knowledge distillation loss for local client objectives collaborating through a global model. We develop an alternating proximal gradient update for solving this compressed personalization problem, and analyze its convergence properties. Numerically, we validate that QuPeD outperforms competing personalized FL methods, FedAvg, and local training of clients in various heterogeneous settings.