On the Benefits of Sampling in Privacy Preserving Statistical Analysis on Distributed Databases

On the Benefits of Sampling in Privacy Preserving Statistical Analysis on Distributed Databases
复制标题

分布式数据库隐私保护统计分析中抽样的好处

DOI:
--
复制
发表时间:
2013
期刊:
arXiv.org
影响因子:
--
通讯作者:
S. Rane
S. Rane
中科院分区:
--
文献类型:
--
作者:
Bing;Ye Wang;S. Rane

文献摘要

被引文献

相似文献

我们认为相互不信任的策展人拥有一个垂直分区数据库的问题,其中包含有关一组个人的信息。目的是使授权方从数据库中获取汇总(统计)信息,同时保护个人的隐私,我们使用差异隐私将其正式化。不受信任的服务器可以促进此过程,该服务器提供存储和处理服务,但不应了解有关数据库的任何信息。这项工作描述了一种数据释放机制,该机制采用随机后(PRAM),加密和随机抽样来维持隐私,同时允许授权方对数据进行准确的统计分析。加密可确保存储服务器没有获得有关数据库的信息,而婴儿车和采样确保对授权方保持个人隐私。我们表征了与单独使用婴儿车相比,随机抽样与婴儿车的组成增加了系统的差异隐私。我们还通过界定估计误差(真正的经验分布与估计分布之间的预期L2 -norm误差)来分析系统的统计效用 - 作为样本数量,下载噪声和其他系统参数的函数。我们的分析表明,增加的下架噪声与减少样品数量以保持所需的隐私水平之间的权衡,我们确定了平衡这种权衡并最大化效用的最佳样品数量。在使用UCI“成人数据集”和合成生成数据的实验模拟中,我们确认理论上预测的最佳样品数量确实达到了接近最小的经验误差,并且我们的分析误差界限与经验结果非常匹配。
We consider a problem where mutually untrusting curators possess portions of a vertically partitioned database containing information about a set of individuals. The goal is to enable an authorized party to obtain aggregate (statistical) information from the database while protecting the privacy of the individuals, which we formalize using Differential Privacy. This process can be facilitated by an untrusted server that provides storage and processing services but should not learn anything about the database. This work describes a data release mechanism that employs Post Randomization (PRAM), encryption and random sampling to maintain privacy, while allowing the authorized party to conduct an accurate statistical analysis of the data. Encryption ensures that the storage server obtains no information about the database, while PRAM and sampling ensures individual privacy is maintained against the authorized party. We characterize how much the composition of random sampling with PRAM increases the differential privacy of system compared to using PRAM alone. We also analyze the statistical utility of our system, by bounding the estimation error - the expected l2-norm error between the true empirical distribution and the estimated distribution - as a function of the number of samples, PRAM noise, and other system parameters. Our analysis shows a tradeoff between increasing PRAM noise versus decreasing the number of samples to maintain a desired level of privacy, and we determine the optimal number of samples that balances this tradeoff and maximizes the utility. In experimental simulations with the UCI "Adult Data Set" and with synthetically generated data, we confirm that the theoretically predicted optimal number of samples indeed achieves close to the minimal empirical error, and that our analytical error bounds match well with the empirical results.