Controlling Privacy Loss in Sampling Schemes: An Analysis of Stratified and Cluster Sampling

Controlling Privacy Loss in Sampling Schemes: An Analysis of Stratified and Cluster Sampling
复制标题

DOI:
10.4230/lipics.forc.2022.1
复制
发表时间:
2020-07
期刊:
--
影响因子:
--
通讯作者:
Mark Bun;Jörg Drechsler;Marco Gaboardi;Audra McMillan;Jayshree Sarathy
Mark Bun;Jörg Drechsler;Marco Gaboardi;Audra McMillan;Jayshree Sarathy
中科院分区:
其他
文献类型:
--
作者:
Mark Bun;Jörg Drechsler;Marco Gaboardi;Audra McMillan;Jayshree Sarathy

文献摘要

被引文献

相似文献

抽样方案是统计学、调查设计和算法设计中的基本工具。差分隐私的一个基本结果是,在人口的简单随机样本上运行的差分隐私机制比在整个人口上运行的相同算法提供更强的隐私保证。然而,在实践中,抽样设计往往比简单的,数据独立的抽样计划,在以前的工作中解决的更复杂。在这项工作中,我们将隐私放大结果的研究扩展到更复杂的数据依赖的采样方案。我们发现,这些抽样方案不仅往往无法放大隐私,实际上还可能导致隐私退化。我们分析了隐私的影响,普遍的整群抽样和分层抽样范式,以及提供一些见解的研究更一般的抽样设计。
Sampling schemes are fundamental tools in statistics, survey design, and algorithm design. A fundamental result in differential privacy is that a differentially private mechanism run on a simple random sample of a population provides stronger privacy guarantees than the same algorithm run on the entire population. However, in practice, sampling designs are often more complex than the simple, data-independent sampling schemes that are addressed in prior work. In this work, we extend the study of privacy amplification results to more complex, data-dependent sampling schemes. We find that not only do these sampling schemes often fail to amplify privacy, they can actually result in privacy degradation. We analyze the privacy implications of the pervasive cluster sampling and stratified sampling paradigms, as well as provide some insight into the study of more general sampling designs.