Calibrate: Frequency Estimation and Heavy Hitter Identification with Local Differential Privacy via Incorporating Prior Knowledge

Calibrate: Frequency Estimation and Heavy Hitter Identification with Local Differential Privacy via Incorporating Prior Knowledge
复制标题

DOI:
10.1109/infocom.2019.8737527
复制
发表时间:
2018-12
期刊:
IEEE INFOCOM 2019 - IEEE Conference on Computer Communications
影响因子:
--
通讯作者:
Jinyuan Jia;N. Gong
Jinyuan Jia;N. Gong
中科院分区:
其他
文献类型:
--
作者:
Jinyuan Jia;N. Gong

文献摘要

被引文献

相似文献

估计人群中某些项目的频率是数据分析的基本步骤,它可以使更多高级数据分析(例如,重型击球手识别,经常模式采矿),客户软件优化以及检测浏览器中用户设置的不必要或恶意劫持。频率估算和重型击球手识别局部差分隐私(LDP)保护用户隐私以及现有的LDP算法。杠杆作用1)关于估计项目频率的噪声和2)关于真实项目频率的知识。知识。我们设计的校准可以通过统计推理加入先验的知识将有关噪声和真实项目频率的先验知识分别为两个概率分布,鉴于两个概率分布和现有的LDP算法产生的项目的估计频率,我们的校准计算了项目频率和项目的条件概率分布使用条件概率分布的含义作为项目的校准频率。稀疏性我们通过统计和机器学习的整合技术来应对挑战。
Estimating frequencies of certain items among a population is a basic step in data analytics, which enables more advanced data analytics (e.g., heavy hitter identification, frequent pattern mining), client software optimization, and detecting unwanted or malicious hijacking of user settings in browsers. Frequency estimation and heavy hitter identification with local differential privacy (LDP) protect user privacy as well as the data collector. Existing LDP algorithms cannot leverage 1) prior knowledge about the noise in the estimated item frequencies and 2) prior knowledge about the true item frequencies. As a result, they achieve suboptimal performance in practice. In this work, we aim to design LDP algorithms that can leverage such prior knowledge. Specifically, we design Calibrate to incorporate the prior knowledge via statistical inference. Calibrate can be appended to an existing LDP algorithm to reduce its estimation errors. We model the prior knowledge about the noise and the true item frequencies as two probability distributions, respectively. Given the two probability distributions and an estimated frequency of an item produced by an existing LDP algorithm, our Calibrate computes the conditional probability distribution of the item’s frequency and uses the mean of the conditional probability distribution as the calibrated frequency for the item. It is challenging to estimate the two probability distributions due to data sparsity. We address the challenge via integrating techniques from statistics and machine learning. Our empirical results on two real-world datasets show that Calibrate significantly outperforms state-of-the-art LDP algorithms for frequency estimation and heavy hitter identification.