Differentially Private Image Classification from Features

Differentially Private Image Classification from Features
复制标题

DOI:
10.48550/arxiv.2211.13403
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Harsh Mehta;Walid Krichene;Abhradeep Thakurta;Alexey Kurakin;Ashok Cutkosky
Harsh Mehta;Walid Krichene;Abhradeep Thakurta;Alexey Kurakin;Ashok Cutkosky
中科院分区:
其他
文献类型:
--
作者:
Harsh Mehta;Walid Krichene;Abhradeep Thakurta;Alexey Kurakin;Ashok Cutkosky

文献摘要

相似文献

利用迁移学习最近被证明是训练具有差分隐私(DP)的大型模型的有效策略。此外,令人惊讶的是,最近的研究发现,仅对预训练模型的最后一层进行私人训练,可以提供DP的最佳效用。虽然过去的研究在很大程度上依赖于像DP-SGD这样的算法来训练大型模型,但在私下从特征中学习的特定情况下,我们观察到计算负担足够低,可以使用更复杂的优化方案,包括二阶方法。为此,我们系统地探讨了损失函数和优化算法等设计参数的影响。我们发现,虽然常用的逻辑回归在非私人环境中的表现优于线性回归,但在私人环境中情况相反。我们发现,线性回归是更有效的比逻辑回归从隐私和计算方面,特别是在更严格的隐私值(1$)。在优化方面,我们也探索使用牛顿的方法,并发现二阶信息是相当有帮助的,即使与隐私,虽然好处显着减少与更严格的隐私保证。虽然这两种方法都使用二阶信息,但最小二乘法在较低的ε下有效,而牛顿法在较大的ε值下有效。为了将两者的优点联合收割机结合起来,我们提出了一种名为DP-FC的新算法,该算法利用特征协方差而不是逻辑回归损失的Hessian,并且在我们尝试的所有$\n $值中表现良好。有了这个,我们在ImageNet-1 k,CIFAR-100和CIFAR-10上获得了新的SOTA结果,涵盖了通常考虑的所有值。最值得注意的是,在ImageNet-1 K上,我们在(8,$8 * 10^{-7}$)-DP下获得了88\%的top-1准确率,在(0.1,$8 * 10^{-7}$)-DP下获得了84.3\%。
Leveraging transfer learning has recently been shown to be an effective strategy for training large models with Differential Privacy (DP). Moreover, somewhat surprisingly, recent works have found that privately training just the last layer of a pre-trained model provides the best utility with DP. While past studies largely rely on algorithms like DP-SGD for training large models, in the specific case of privately learning from features, we observe that computational burden is low enough to allow for more sophisticated optimization schemes, including second-order methods. To that end, we systematically explore the effect of design parameters such as loss function and optimization algorithm. We find that, while commonly used logistic regression performs better than linear regression in the non-private setting, the situation is reversed in the private setting. We find that linear regression is much more effective than logistic regression from both privacy and computational aspects, especially at stricter epsilon values ($\epsilon<1$). On the optimization side, we also explore using Newton's method, and find that second-order information is quite helpful even with privacy, although the benefit significantly diminishes with stricter privacy guarantees. While both methods use second-order information, least squares is effective at lower epsilons while Newton's method is effective at larger epsilon values. To combine the benefits of both, we propose a novel algorithm called DP-FC, which leverages feature covariance instead of the Hessian of the logistic regression loss and performs well across all $\epsilon$ values we tried. With this, we obtain new SOTA results on ImageNet-1k, CIFAR-100 and CIFAR-10 across all values of $\epsilon$ typically considered. Most remarkably, on ImageNet-1K, we obtain top-1 accuracy of 88\% under (8, $8 * 10^{-7}$)-DP and 84.3\% under (0.1, $8 * 10^{-7}$)-DP.