Personalized Differentially Private Federated Learning without Exposing Privacy Budgets

Personalized Differentially Private Federated Learning without Exposing Privacy Budgets
复制标题

DOI:
10.1145/3583780.3615247
复制
发表时间:
2023-10
期刊:
Proceedings of the 32nd ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Junxu Liu;Jian Lou;Li Xiong;Xiaofeng Meng
Junxu Liu;Jian Lou;Li Xiong;Xiaofeng Meng
中科院分区:
其他
文献类型:
--
作者:
Junxu Liu;Jian Lou;Li Xiong;Xiaofeng Meng

文献摘要

相似文献

跨竖井联邦学习(FL)的迅速崛起是由于它能够减轻协作训练期间的数据泄露。考虑到不同客户端的不同隐私需求,为了进一步提供严格的隐私保护,提出了一种个性化差异隐私联邦学习(PDP-FL)的隐私增强工作。但是,PDP-FL[20]的现有解决方案假定所有客户机的原始隐私预算都应该由服务器收集。然后,通过促进隐私首选项分区(即,将所有客户端划分为多个隐私组),直接利用这些值来改进模型实用程序。然而,这是不现实的,因为原始的隐私预算可以提供相当多的信息和敏感。在这项工作中,我们的目标是通过仅基于客户端的噪声模型更新间接划分隐私偏好来实现PDP-FL,而不会暴露客户端的原始隐私预算。问题的关键在于,有噪声的更新可能受到DP噪声和非iid客户端数据两个纠缠因素的影响,是否有可能通过解开这两个影响因素来揭示隐私偏好是未知的。为了克服这一障碍,我们系统地研究了在什么条件下客户端的模型更新主要受噪声水平而不是数据分布的影响这一尚未探索的问题。然后,我们提出了一种简单而有效的基于噪声更新L2范数聚类的策略,该策略可以集成到普通PDP-FL中以保持相同的性能。实验结果证明了该算法的有效性和可行性。
The meteoric rise of cross-silo Federated Learning (FL) is due to its ability to mitigate data breaches during collaborative training. To further provide rigorous privacy protection with consideration of the varying privacy requirements across different clients, a privacy-enhanced line of work on personalized differentially private federated learning (PDP-FL) has been proposed. However, the existing solution for PDP-FL [20] assumes the raw privacy budgets of all clients should be collected by the server. These values are then directly utilized to improve the model utility via facilitating the privacy preferences partitioning (i.e., partitioning all clients into multiple privacy groups). It is however non-realistic because the raw privacy budgets can be quite informative and sensitive. In this work, our goal is to achieve PDP-FL without exposing clients' raw privacy budgets by indirectly partitioning the privacy preferences solely based on clients' noisy model updates. The crux lies in the fact that the noisy updates could be influenced by two entangled factors of DP noises and non-IID clients' data, leaving it unknown whether it is possible to uncover privacy preferences by disentangling the two affecting factors. To overcome the hurdle, we systematically investigate the unexplored question of under what conditions can the model updates of clients be primarily influenced by noise levels rather than data distribution. Then, we propose a simple yet effective strategy based on clustering the L2 norm of the noisy updates, which can be integrated into the vanilla PDP-FL to maintain the same performance. Experimental results demonstrate the effectiveness and feasibility of our privacy-budget-agnostic PDP-FL method.