Renyi Differential Privacy of the Subsampled Shuffle Model in Distributed Learning

Renyi Differential Privacy of the Subsampled Shuffle Model in Distributed Learning
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
Antonious M. Girgis;Deepesh Data;S. Diggavi
Antonious M. Girgis;Deepesh Data;S. Diggavi
中科院分区:
其他
文献类型:
--
作者:
Antonious M. Girgis;Deepesh Data;S. Diggavi

文献摘要

相似文献

我们在分布式学习框架中研究隐私,其中客户端通过与我们需要隐私的服务器交互来协作构建学习模型。在随机优化和联邦学习(FL)范式的推动下,我们关注每轮中随机对一小部分数据样本进行子采样以参与学习过程的情况,这也使得隐私放大成为可能。为了获得更强的本地隐私保证,我们在洗牌隐私模型中进行研究,其中每个客户端使用本地差分隐私(LDP)机制随机化其响应,并且服务器仅接收客户端响应的随机排列(洗牌),而不将其与每个客户端关联。本文的主要结果是在该子采样洗牌隐私模型中对离散随机化机制进行隐私优化性能权衡。这是通过一种新的理论技术来分析子采样洗牌模型的 Renyi 差分隐私 (RDP) 来实现的。我们通过数值证明,对于重要的制度,与子采样混洗模型的最先进的近似差分隐私(DP)保证(具有强组合)相比,我们的界限通过组合在隐私保证方面产生了显着的改进。我们还使用真实数据集证明了隐私学习性能操作点在数值上的显着改进。
We study privacy in a distributed learning framework, where clients collaboratively build a learning model iteratively through interactions with a server from whom we need privacy. Motivated by stochastic optimization and the federated learning (FL) paradigm, we focus on the case where a small fraction of data samples are randomly sub-sampled in each round to participate in the learning process, which also enables privacy amplification. To obtain even stronger local privacy guarantees, we study this in the shuffle privacy model, where each client randomizes its response using a local differentially private (LDP) mechanism and the server only receives a random permutation (shuffle) of the clients' responses without their association to each client. The principal result of this paper is a privacy-optimization performance trade-off for discrete randomization mechanisms in this sub-sampled shuffle privacy model. This is enabled through a new theoretical technique to analyze the Renyi Differential Privacy (RDP) of the sub-sampled shuffle model. We numerically demonstrate that, for important regimes, with composition our bound yields significant improvement in privacy guarantee over the state-of-the-art approximate Differential Privacy (DP) guarantee (with strong composition) for sub-sampled shuffled models. We also demonstrate numerically significant improvement in privacy-learning performance operating point using real data sets.