Benchmarking Private Population Data Release Mechanisms: Synthetic Data vs. TopDown

Benchmarking Private Population Data Release Mechanisms: Synthetic Data vs. TopDown
复制标题

DOI:
10.48550/arxiv.2401.18024
复制
发表时间:
2024-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Aadyaa Maddi;Swadhin Routray;Alexander Goldberg;Giulia Fanti
Aadyaa Maddi;Swadhin Routray;Alexander Goldberg;Giulia Fanti
中科院分区:
其他
文献类型:
--
作者:
Aadyaa Maddi;Swadhin Routray;Alexander Goldberg;Giulia Fanti

文献摘要

相似文献

差分隐私(DP)越来越多地用于保护分层的、表格式的人口数据(如人口普查数据)的发布。在这种情况下实现DP的一种常见方法是释放对预定义查询集的噪声响应。例如,这是美国人口普查局使用的TopDown算法的方法。这样的方法有一个重要的缺点:它们不能回答它们没有优化的查询。一个有吸引力的替代方案是生成DP合成数据,该数据从一些生成分布中提取。与TopDown方法一样,合成数据也可以优化以回答特定查询,同时还允许数据用户稍后提交对合成人口数据的任意查询。据我们所知,还没有一个头对头的实证比较这些方法。本研究对TopDown算法和私有合成数据生成进行了比较,以确定查询复杂性、分布内与分布外查询以及隐私保证对准确性的影响。我们的研究结果表明,在分布查询,TopDown算法实现了显着更好的隐私保真度的权衡比我们评估的任何合成数据的方法,例如,在我们的实验中,TopDown实现了至少$20\times$较低的错误计数查询比领先的合成数据方法在相同的隐私预算。我们的研究结果为从业者和合成数据研究社区提供了指导方针。
Differential privacy (DP) is increasingly used to protect the release of hierarchical, tabular population data, such as census data. A common approach for implementing DP in this setting is to release noisy responses to a predefined set of queries. For example, this is the approach of the TopDown algorithm used by the US Census Bureau. Such methods have an important shortcoming: they cannot answer queries for which they were not optimized. An appealing alternative is to generate DP synthetic data, which is drawn from some generating distribution. Like the TopDown method, synthetic data can also be optimized to answer specific queries, while also allowing the data user to later submit arbitrary queries over the synthetic population data. To our knowledge, there has not been a head-to-head empirical comparison of these approaches. This study conducts such a comparison between the TopDown algorithm and private synthetic data generation to determine how accuracy is affected by query complexity, in-distribution vs. out-of-distribution queries, and privacy guarantees. Our results show that for in-distribution queries, the TopDown algorithm achieves significantly better privacy-fidelity tradeoffs than any of the synthetic data methods we evaluated; for instance, in our experiments, TopDown achieved at least $20\times$ lower error on counting queries than the leading synthetic data method at the same privacy budget. Our findings suggest guidelines for practitioners and the synthetic data research community.