Toward Distribution Estimation under Local Differential Privacy with Small Samples

Toward Distribution Estimation under Local Differential Privacy with Small Samples
复制标题

DOI:
10.1515/popets-2018-0022
复制
发表时间:
2018-06
影响因子:
--
通讯作者:
Takao Murakami;H. Hino;Jun Sakuma
Takao Murakami;H. Hino;Jun Sakuma
中科院分区:
--
文献类型:
--
作者:
Takao Murakami;H. Hino;Jun Sakuma

文献摘要

被引文献

相似文献

摘要最近对局部模型中的离散分布估计进行了一些研究,其中用户自己混淆他们的个人数据(例如,调查中的位置、响应),数据收集器从混淆的数据中估计原始个人数据的分布。与集中式模式不同,在集中式模式中,受信任的数据库管理员可以访问所有用户的个人数据,而本地模式不存在数据泄露的风险。该模型中一个具有代表性的隐私度量是局部差异隐私,它通过一个称为隐私预算的参数∈来控制信息泄漏量。当∈很小时,会给个人数据增加大量的噪音,因此用户的隐私得到了强有力的保护。然而,当用户ℕ的数量较小时(例如,小型企业可能无法收集大样本),或者当大多数用户采用较小的∈值时,估计分布变得非常具有挑战性。本文的目的是准确估计上述情况下的分布。为了实现这一目标,我们重点研究了一种最新的统计推理方法EM(期望最大化)重构方法,并利用Rilstone等人的理论提出了一种修正其估计误差(即估计与真实值之间的差异)的方法。在一定的假设条件下,证明了该方法降低了均方误差,并用三个大规模数据集对该方法进行了评估,其中两个数据集包含位置数据,另一个数据集包含人口普查数据。结果表明,当ℕ或∈较小时,该方法在所有数据集上的性能都明显优于EM重建方法。
Abstract A number of studies have recently been made on discrete distribution estimation in the local model, in which users obfuscate their personal data (e.g., location, response in a survey) by themselves and a data collector estimates a distribution of the original personal data from the obfuscated data. Unlike the centralized model, in which a trusted database administrator can access all users’ personal data, the local model does not suffer from the risk of data leakage. A representative privacy metric in this model is LDP (Local Differential Privacy), which controls the amount of information leakage by a parameter ∈ called privacy budget. When ∈ is small, a large amount of noise is added to the personal data, and therefore users’ privacy is strongly protected. However, when the number of users ℕ is small (e.g., a small-scale enterprise may not be able to collect large samples) or when most users adopt a small value of ∈, the estimation of the distribution becomes a very challenging task. The goal of this paper is to accurately estimate the distribution in the cases explained above. To achieve this goal, we focus on the EM (Expectation-Maximization) reconstruction method, which is a state-of-the-art statistical inference method, and propose a method to correct its estimation error (i.e., difference between the estimate and the true value) using the theory of Rilstone et al. We prove that the proposed method reduces the MSE (Mean Square Error) under some assumptions.We also evaluate the proposed method using three largescale datasets, two of which contain location data while the other contains census data. The results show that the proposed method significantly outperforms the EM reconstruction method in all of the datasets when ℕ or ∈ is small.