Concentrated Differentially Private Federated Learning With Performance Analysis

Concentrated Differentially Private Federated Learning With Performance Analysis
复制标题

DOI:
10.1109/ojcs.2021.3099108
复制
发表时间:
2021
影响因子:
5.9
通讯作者:
Rui Hu;Yuanxiong Guo;Yanmin Gong
Rui Hu;Yuanxiong Guo;Yanmin Gong
中科院分区:
--
文献类型:
--
作者:
Rui Hu;Yuanxiong Guo;Yanmin Gong

文献摘要

相似文献

联合学习利用一组边缘设备来协作训练通用模型,而无需共享其本地数据,并且在用户隐私方面优于传统的基于云的学习方法。然而,最近的模型反演攻击和成员推断攻击表明,在交互式训练过程中共享的模型更新仍然可能泄漏敏感的用户信息。因此,需要在联邦学习中提供严格的差分隐私(DP)保证。提供DP的主要挑战是在DP机制反复引入随机性的情况下保持联邦学习模型的高效用,特别是当服务器不完全可信时。在本文中,我们研究如何提供DP最广泛采用的联邦学习计划,联邦平均。我们的方法结合了局部梯度扰动,安全聚合和零集中差分隐私(zCDP)更好的效用和隐私保护,而无需可信服务器。我们共同考虑的DP机制,客户端采样和数据二次采样在我们的方法中引入的随机性的性能影响,并从理论上分析了收敛速度和端到端的DP保证与非凸损失函数。我们还证明了我们提出的方法具有良好的实用性,隐私权衡通过广泛的数值实验在现实世界的数据集。
Federated learning engages a set of edge devices to collaboratively train a common model without sharing their local data and has advantage in user privacy over traditional cloud-based learning approaches. However, recent model inversion attacks and membership inference attacks have demonstrated that shared model updates during the interactive training process could still leak sensitive user information. Thus, it is desirable to provide rigorous differential privacy (DP) guarantee in federated learning. The main challenge to providing DP is to maintain high utility of federated learning model with repeatedly introduced randomness of DP mechanisms, especially when the server is not fully trusted. In this paper, we investigate how to provide DP to the most widely adopted federated learning scheme, federated averaging. Our approach combines local gradient perturbation, secure aggregation, and zero-concentrated differential privacy (zCDP) for better utility and privacy protection without a trusted server. We jointly consider the performance impacts of randomnesses introduced by the DP mechanism, client sampling and data subsampling in our approach, and theoretically analyze the convergence rate and end-to-end DP guarantee with non-convex loss functions. We also demonstrate that our proposed method has good utility-privacy trade-off through extensive numerical experiments on the real-world dataset.