HePCo: Data-Free Heterogeneous Prompt Consolidation for Continual Federated Learning

HePCo: Data-Free Heterogeneous Prompt Consolidation for Continual Federated Learning
复制标题

DOI:
10.48550/arxiv.2306.09970
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Shaunak Halbe;James Smith;Junjiao Tian;Z. Kira
Shaunak Halbe;James Smith;Junjiao Tian;Z. Kira
中科院分区:
其他
文献类型:
--
作者:
Shaunak Halbe;James Smith;Junjiao Tian;Z. Kira

文献摘要

被引文献

相似文献

在本文中,我们关注持续联邦学习(CFL)这一重要但研究不足的问题。在该问题中,服务器与一组客户端进行通信,以便随时间逐步学习新概念,且不共享或存储任何数据。从持续学习和联邦学习两个角度来看,该问题的复杂性因各种挑战而加剧。具体而言,在CFL设置中训练的模型会遭受灾难性遗忘,而客户端之间的数据异质性会使这种情况更加严重。针对该问题的现有尝试往往会给客户端和通信通道带来巨大开销,或者需要访问存储的数据,由于隐私问题,这使得它们不适合在现实世界中使用。在本文中,我们试图在最小化开销成本且不需要访问任何存储数据的情况下解决遗忘和异质性问题。我们通过利用一种基于提示的方法(这样只需传递提示和分类器头部),并提出一种新颖的轻量级生成和蒸馏方案来在服务器上整合客户端模型来实现这一目标。我们针对图像分类阐述了这个问题,并建立了强大的基线用于比较,在CIFAR - 100以及具有挑战性的大规模数据集(如ImageNet - R和DomainNet)上进行了实验。我们的方法比现有方法和我们自己的基线性能高出多达7%,同时显著降低了通信和客户端级别的计算成本。
In this paper, we focus on the important yet understudied problem of Continual Federated Learning (CFL), where a server communicates with a set of clients to incrementally learn new concepts over time without sharing or storing any data. The complexity of this problem is compounded by challenges from both the Continual and Federated Learning perspectives. Specifically, models trained in a CFL setup suffer from catastrophic forgetting which is exacerbated by data heterogeneity across clients. Existing attempts at this problem tend to impose large overheads on clients and communication channels or require access to stored data which renders them unsuitable for real-world use due to privacy. In this paper, we attempt to tackle forgetting and heterogeneity while minimizing overhead costs and without requiring access to any stored data. We achieve this by leveraging a prompting based approach (such that only prompts and classifier heads have to be communicated) and proposing a novel and lightweight generation and distillation scheme to consolidate client models at the server. We formulate this problem for image classification and establish strong baselines for comparison, conduct experiments on CIFAR-100 as well as challenging, large-scale datasets like ImageNet-R and DomainNet. Our approach outperforms both existing methods and our own baselines by as much as 7% while significantly reducing communication and client-level computation costs.