Generalized Federated Learning via Sharpness Aware Minimization

Generalized Federated Learning via Sharpness Aware Minimization
复制标题

DOI:
10.48550/arxiv.2206.02618
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Zhe Qu;Xingyu Li;Rui Duan;Yaojiang Liu;Bo Tang;Zhuo Lu
Zhe Qu;Xingyu Li;Rui Duan;Yaojiang Liu;Bo Tang;Zhuo Lu
中科院分区:
其他
文献类型:
--
作者:
Zhe Qu;Xingyu Li;Rui Duan;Yaojiang Liu;Bo Tang;Zhuo Lu

文献摘要

被引文献

相似文献

联邦学习(FL)是一个很有前途的框架,用于对一组客户端执行隐私保护和分布式学习。然而,客户端之间的数据分布往往表现为非iid,即分布移位,这使得高效优化变得困难。为了解决这个问题,许多FL算法专注于通过提高全局模型的性能来减轻客户机之间数据异构的影响。然而,几乎所有的算法都是利用经验风险最小化(Empirical Risk Minimization, ERM)作为局部优化器,这很容易使全局模型陷入陡谷,增加局部客户端部分的较大偏差。因此,在本文中,我们重新审视了FL中分布移位问题的解决方案,重点关注局部学习的通用性。为此,我们提出了一种基于锐利感知最小化(锐利感知最小化)局部优化器的通用有效算法\texttt{FedSAM},并开发了一种用于连接局部模型和全局模型的动量FL算法\texttt{MoFedSAM}。从理论上分析了这两种算法的收敛性,并给出了\texttt{FedSAM}的泛化界。从经验上看,我们提出的算法大大优于现有的FL研究,并显著降低了学习偏差。
Federated Learning (FL) is a promising framework for performing privacy-preserving, distributed learning with a set of clients. However, the data distribution among clients often exhibits non-IID, i.e., distribution shift, which makes efficient optimization difficult. To tackle this problem, many FL algorithms focus on mitigating the effects of data heterogeneity across clients by increasing the performance of the global model. However, almost all algorithms leverage Empirical Risk Minimization (ERM) to be the local optimizer, which is easy to make the global model fall into a sharp valley and increase a large deviation of parts of local clients. Therefore, in this paper, we revisit the solutions to the distribution shift problem in FL with a focus on local learning generality. To this end, we propose a general, effective algorithm, \texttt{FedSAM}, based on Sharpness Aware Minimization (SAM) local optimizer, and develop a momentum FL algorithm to bridge local and global models, \texttt{MoFedSAM}. Theoretically, we show the convergence analysis of these two algorithms and demonstrate the generalization bound of \texttt{FedSAM}. Empirically, our proposed algorithms substantially outperform existing FL studies and significantly decrease the learning deviation.