Stolen Memories: Leveraging Model Memorization for Calibrated White-Box Membership Inference

Stolen Memories: Leveraging Model Memorization for Calibrated White-Box Membership Inference
复制标题

DOI:
--
复制
发表时间:
2019-06
期刊:
Proceedings of the 16th ACM International Conference on Computing Frontiers
影响因子:
--
通讯作者:
Klas Leino;Matt Fredrikson
Klas Leino;Matt Fredrikson
中科院分区:
其他
文献类型:
--
作者:
Klas Leino;Matt Fredrikson

文献摘要

被引文献

相似文献

隶属推理(MI)攻击利用了机器学习算法有时会通过学习模型泄露训练数据的信息这一事实。在这项工作中,我们研究了白盒设置中的隶属推理,以利用模型的内部,这在以前的工作中没有得到有效利用。利用关于深度神经网络中过度拟合如何发生的新见解,我们展示了模型对特征的特殊使用如何为白盒攻击者提供成员资格的证据-即使模型的黑箱行为似乎泛化得很好-并证明这种攻击优于先前的黑箱方法。考虑到有效的攻击应该具有提供自信的正向推理的能力,我们发现以前的攻击通常不能为自信地推断隶属度提供有意义的基础,而我们的攻击可以有效地校准为高精度。最后,我们研究了针对MI攻击的常用防御措施,发现(1)较小的泛化误差不足以防止对真实模型的攻击,(2)虽然较小的差分隐私降低了攻击的有效性,但这通常会对模型的准确性造成重大损失;对于有时在实践中使用的较大的$\epsilon$(例如,$\epsilon=16$),攻击可以达到与未受保护的模型几乎相同的精度。
Membership inference (MI) attacks exploit the fact that machine learning algorithms sometimes leak information about their training data through the learned model. In this work, we study membership inference in the white-box setting in order to exploit the internals of a model, which have not been effectively utilized by previous work. Leveraging new insights about how overfitting occurs in deep neural networks, we show how a model's idiosyncratic use of features can provide evidence for membership to white-box attackers---even when the model's black-box behavior appears to generalize well---and demonstrate that this attack outperforms prior black-box methods. Taking the position that an effective attack should have the ability to provide confident positive inferences, we find that previous attacks do not often provide a meaningful basis for confidently inferring membership, whereas our attack can be effectively calibrated for high precision. Finally, we examine popular defenses against MI attacks, finding that (1) smaller generalization error is not sufficient to prevent attacks on real models, and (2) while small-$\epsilon$-differential privacy reduces the attack's effectiveness, this often comes at a significant cost to the model's accuracy; and for larger $\epsilon$ that are sometimes used in practice (e.g., $\epsilon=16$), the attack can achieve nearly the same accuracy as on the unprotected model.