Orchestra: Unsupervised Federated Learning via Globally Consistent Clustering

Orchestra: Unsupervised Federated Learning via Globally Consistent Clustering
复制标题

DOI:
10.48550/arxiv.2205.11506
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Ekdeep Singh Lubana;Chi Ian Tang;F. Kawsar;R. Dick;Akhil Mathur
Ekdeep Singh Lubana;Chi Ian Tang;F. Kawsar;R. Dick;Akhil Mathur
中科院分区:
其他
文献类型:
--
作者:
Ekdeep Singh Lubana;Chi Ian Tang;F. Kawsar;R. Dick;Akhil Mathur

文献摘要

相似文献

联合学习通常用于标签容易获得的任务(例如,下一词预测)。放松这一限制需要设计非监督学习技术来支持联合训练所需的特性:对统计/系统异构性的健壮性、参与者数量的可扩展性和通信效率。以前在这个主题上的工作集中在直接扩展集中式自我监督学习技术,这些技术没有设计成具有上面列出的属性。为了解决这种情况,我们提出了一种新的无监督联合学习技术Orchestra,它利用联邦的层次结构来协调分布式聚类任务,并将客户数据全局一致地划分为可区分的簇。我们展示了Orchestra中的算法流水线在线性探测下保证了良好的泛化性能,使其在包括异构性、客户端数量、参与率和局部历元的变化在内的广泛条件下的表现优于其他技术。
Federated learning is generally used in tasks where labels are readily available (e.g., next word prediction). Relaxing this constraint requires design of unsupervised learning techniques that can support desirable properties for federated training: robustness to statistical/systems heterogeneity, scalability with number of participants, and communication efficiency. Prior work on this topic has focused on directly extending centralized self-supervised learning techniques, which are not designed to have the properties listed above. To address this situation, we propose Orchestra, a novel unsupervised federated learning technique that exploits the federation's hierarchy to orchestrate a distributed clustering task and enforce a globally consistent partitioning of clients' data into discriminable clusters. We show the algorithmic pipeline in Orchestra guarantees good generalization performance under a linear probe, allowing it to outperform alternative techniques in a broad range of conditions, including variation in heterogeneity, number of clients, participation ratio, and local epochs.