Arbitrary Decisions are a Hidden Cost of Differentially Private Training

Arbitrary Decisions are a Hidden Cost of Differentially Private Training
复制标题

DOI:
10.1145/3593013.3594103
复制
发表时间:
2023-02
期刊:
Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency
影响因子:
--
通讯作者:
B. Kulynych;Hsiang Hsu;C. Troncoso;F. Calmon
B. Kulynych;Hsiang Hsu;C. Troncoso;F. Calmon
中科院分区:
其他
文献类型:
--
作者:
B. Kulynych;Hsiang Hsu;C. Troncoso;F. Calmon

文献摘要

相似文献

用于隐私保护的机器学习机制通常旨在保证模型训练期间的差异隐私(DP)。实用的DP确保训练方法在将模型参数与隐私敏感数据进行匹配时使用随机化(例如,将高斯噪声添加到修剪的梯度中)。我们证明了这种随机化会导致预测的多重性:对于给定的输入示例,由同等私有模型预测的输出取决于训练中使用的随机性。因此,对于给定的输入,如果重新训练模型,即使使用相同的训练数据集,预测输出也会有很大的变化。发展规划培训的预测多样性成本尚未得到研究,目前既没有对模型设计者和利益攸关方进行审计,也没有向他们通报。我们得出了可靠地估计预测多重性所需的重新训练次数的界限。我们从理论上和大量的实验中分析了三种DP保证算法:输出扰动、目标扰动和DP-SGD的预测重数代价。我们证明了预测的多样性程度随着隐私级别的增加而增加,并且在数据中的个人和人口群体之间分布不均匀。由于在训练期间用于确保DP的随机性解释了一些例子的预测,我们的结果突显了在高风险环境中由不同私有模型支持的决策的正当性面临的根本挑战。我们的结论是,实践者在将其DP确保算法部署到具有个人级别后果的应用程序之前,应该对其预测的多重性进行审计。
Mechanisms used in privacy-preserving machine learning often aim to guarantee differential privacy (DP) during model training. Practical DP-ensuring training methods use randomization when fitting model parameters to privacy-sensitive data (e.g., adding Gaussian noise to clipped gradients). We demonstrate that such randomization incurs predictive multiplicity: for a given input example, the output predicted by equally-private models depends on the randomness used in training. Thus, for a given input, the predicted output can vary drastically if a model is re-trained, even if the same training dataset is used. The predictive-multiplicity cost of DP training has not been studied, and is currently neither audited for nor communicated to model designers and stakeholders. We derive a bound on the number of re-trainings required to estimate predictive multiplicity reliably. We analyze—both theoretically and through extensive experiments—the predictive-multiplicity cost of three DP-ensuring algorithms: output perturbation, objective perturbation, and DP-SGD. We demonstrate that the degree of predictive multiplicity rises as the level of privacy increases, and is unevenly distributed across individuals and demographic groups in the data. Because randomness used to ensure DP during training explains predictions for some examples, our results highlight a fundamental challenge to the justifiability of decisions supported by differentially-private models in high-stakes settings. We conclude that practitioners should audit the predictive multiplicity of their DP-ensuring algorithms before deploying them in applications of individual-level consequence.