Mean Estimation with User-level Privacy under Data Heterogeneity

Mean Estimation with User-level Privacy under Data Heterogeneity
复制标题

DOI:
10.48550/arxiv.2307.15835
复制
发表时间:
2023-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Rachel Cummings;V. Feldman;Audra McMillan;Kunal Talwar
Rachel Cummings;V. Feldman;Audra McMillan;Kunal Talwar
中科院分区:
其他
文献类型:
--
作者:
Rachel Cummings;V. Feldman;Audra McMillan;Kunal Talwar

文献摘要

被引文献

相似文献

许多现代数据分析任务中的一个关键挑战是用户数据是异构的。不同的用户可能拥有截然不同数量的数据点。更重要的是,不能假设所有用户都从相同的基础分布中采样。确实如此,例如在语言数据中,不同的语音风格会导致数据异质性。在这项工作中,我们提出了一种简单的异构用户数据模型,该模型允许用户数据在数据分布和数据数量上都不同,并提供了一种在保留用户级差异隐私的同时估计总体水平均值的方法。我们证明了估计器的渐近最优性,并证明了在我们引入的设置中可实现的误差的一般下界。
A key challenge in many modern data analysis tasks is that user data are heterogeneous. Different users may possess vastly different numbers of data points. More importantly, it cannot be assumed that all users sample from the same underlying distribution. This is true, for example in language data, where different speech styles result in data heterogeneity. In this work we propose a simple model of heterogeneous user data that allows user data to differ in both distribution and quantity of data, and provide a method for estimating the population-level mean while preserving user-level differential privacy. We demonstrate asymptotic optimality of our estimator and also prove general lower bounds on the error achievable in the setting we introduce.