FLAME: Differentially Private Federated Learning in the Shuffle Model

FLAME: Differentially Private Federated Learning in the Shuffle Model
复制标题

DOI:
10.1609/aaai.v35i10.17053
复制
发表时间:
2020-09
期刊:
--
影响因子:
--
通讯作者:
Ruixuan Liu;Yang Cao;Hong Chen;Ruoyang Guo;Masatoshi Yoshikawa
Ruixuan Liu;Yang Cao;Hong Chen;Ruoyang Guo;Masatoshi Yoshikawa
中科院分区:
其他
文献类型:
--
作者:
Ruixuan Liu;Yang Cao;Hong Chen;Ruoyang Guo;Masatoshi Yoshikawa

文献摘要

被引文献

相似文献

联邦学习(FL)是一种很有前途的机器学习范式,它使分析器能够在不收集用户原始数据的情况下训练模型。为了保证用户的隐私,差分私有联邦学习得到了广泛的研究。现有的工作主要是基于策展人模型或局部模型的差异隐私。然而,这两种方法都有优点和缺点。策展人模型允许更高的准确性,但需要一个值得信赖的分析器。在本地模型中,用户在将本地数据发送到分析器之前将其随机化,不需要可信的分析器,但精度有限。在这项工作中,通过利用最近提出的差异隐私洗牌模型中的\textit{privacy amplification}效应,我们实现了两个世界中的最佳,即,馆长模型的准确性和强大的隐私性,而不依赖于任何可信方。我们首先提出了shuffle模型中的FL框架和从现有工作扩展的简单协议(SS-Simple)。我们发现,SS-Simple只提供了一个不足的隐私放大效果在FL,因为模型参数的尺寸是相当大的。为了解决这个问题,我们提出了一个增强的协议(SS-Double),通过子采样来增加隐私放大效果。此外,当模型大小大于用户群体时,为了提高效用,我们提出了一种具有梯度稀疏化技术的高级协议(SS-Topk)。我们还提供了理论分析和数值评估的隐私放大所提出的协议。在真实数据集上的实验表明,SS-Topk比基于局部模型的FL提高了60.7%的测试准确率,比基于策展人模型的FL提高了33.94%。与非私有FL相比,我们的协议SS-Topk在(2.348,5e-6)-DP下每一个epoch仅损失1.48%的准确性。
Federated Learning (FL) is a promising machine learning paradigm that enables the analyzer to train a model without collecting users' raw data. To ensure users' privacy, differentially private federated learning has been intensively studied. The existing works are mainly based on the curator model or local model of differential privacy. However, both of them have pros and cons. The curator model allows greater accuracy but requires a trusted analyzer. In the local model where users randomize local data before sending them to the analyzer, a trusted analyzer is not required but the accuracy is limited. In this work, by leveraging the \textit{privacy amplification} effect in the recently proposed shuffle model of differential privacy, we achieve the best of two worlds, i.e., accuracy in the curator model and strong privacy without relying on any trusted party. We first propose an FL framework in the shuffle model and a simple protocol (SS-Simple) extended from existing work. We find that SS-Simple only provides an insufficient privacy amplification effect in FL since the dimension of the model parameter is quite large. To solve this challenge, we propose an enhanced protocol (SS-Double) to increase the privacy amplification effect by subsampling. Furthermore, for boosting the utility when the model size is greater than the user population, we propose an advanced protocol (SS-Topk) with gradient sparsification techniques. We also provide theoretical analysis and numerical evaluations of the privacy amplification of the proposed protocols. Experiments on real-world dataset validate that SS-Topk improves the testing accuracy by 60.7% than the local model based FL. We highlight an observation that SS-Topk improves the accuracy by 33.94\% than the curator model based FL without any trusted party. Compared with non-private FL, our protocol SS-Topk only lose 1.48% accuracy under (2.348, 5e-6)-DP per epoch.