Covariance-Aware Private Mean Estimation Without Private Covariance Estimation

Covariance-Aware Private Mean Estimation Without Private Covariance Estimation
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Gavin Brown;Marco Gaboardi;Adam D. Smith;Jonathan Ullman;Lydia Zakynthinou
Gavin Brown;Marco Gaboardi;Adam D. Smith;Jonathan Ullman;Lydia Zakynthinou
中科院分区:
其他
文献类型:
--
作者:
Gavin Brown;Marco Gaboardi;Adam D. Smith;Jonathan Ullman;Lydia Zakynthinou

文献摘要

相似文献

我们提出了两个具有未知协方差的$ D $维(次)高斯分布的样品差异私有平均值估计器。非正式地,给定的$ n \ gtrsim d/\ alpha^2 $样本来自该分布的均值$ \ mu $和协方差$ \ sigma $,我们的估计器输出$ \ tilde \ mu $,以便$ \ | \ tilde \ mu - \ mu \ | _ {\ sigma} \ leq \ alpha $,其中$ \ | \ cdot \ | _ {\ sigma} $是Mahalanobis距离。具有相同保证的所有先前估计器要么需要在协方差矩阵上具有强大的先验界限,要么需要$ \ omega(d^{3/2})$样本。我们的每个估计器都是基于设计差异化机制的简单,一般的方法,但采用了新的技术步骤,以使估计量私有和样本效率高。我们的第一个估计器使用指数机制示例具有大约最大Tukey深度的点,但仅限于大型Tukey深度的点集。即使对于具有少量对抗性腐败的数据集,其准确性也可以保证。证明这种机制是私人的,需要进行新的分析。我们的第二个估计器将数据集的经验平均值与经验协方差校准的噪声相关,而无需释放协方差本身。它的样本复杂性保证更普遍地适用于Subgaussian分布,尽管对隐私参数的依赖性略有差。对于这两个估计器,都需要仔细的数据预处理以满足差异隐私。
We present two sample-efficient differentially private mean estimators for $d$-dimensional (sub)Gaussian distributions with unknown covariance. Informally, given $n \gtrsim d/\alpha^2$ samples from such a distribution with mean $\mu$ and covariance $\Sigma$, our estimators output $\tilde\mu$ such that $\| \tilde\mu - \mu \|_{\Sigma} \leq \alpha$, where $\| \cdot \|_{\Sigma}$ is the Mahalanobis distance. All previous estimators with the same guarantee either require strong a priori bounds on the covariance matrix or require $\Omega(d^{3/2})$ samples. Each of our estimators is based on a simple, general approach to designing differentially private mechanisms, but with novel technical steps to make the estimator private and sample-efficient. Our first estimator samples a point with approximately maximum Tukey depth using the exponential mechanism, but restricted to the set of points of large Tukey depth. Its accuracy guarantees hold even for data sets that have a small amount of adversarial corruption. Proving that this mechanism is private requires a novel analysis. Our second estimator perturbs the empirical mean of the data set with noise calibrated to the empirical covariance, without releasing the covariance itself. Its sample complexity guarantees hold more generally for subgaussian distributions, albeit with a slightly worse dependence on the privacy parameter. For both estimators, careful preprocessing of the data is required to satisfy differential privacy.