On the Difficulty of Membership Inference Attacks

On the Difficulty of Membership Inference Attacks
复制标题

DOI:
10.1109/cvpr46437.2021.00780
复制
发表时间:
2021-06
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Shahbaz Rezaei;Xin Liu
Shahbaz Rezaei;Xin Liu
中科院分区:
其他
文献类型:
--
作者:
Shahbaz Rezaei;Xin Liu

文献摘要

相似文献

最近的研究提出了对深度模型的成员推理 (MI) 攻击,其目标是推断样本是否已在训练过程中使用。尽管取得了明显的成功,但这些研究仅报告了阳性类别(成员类别)的准确性、精确度和召回率。因此,这些攻击的表现尚未在负类(非成员类)上明确报告。在本文中,我们表明,MI 攻击性能的报告方式常常具有误导性,因为它们存在尚未报告的高误报率或误报率 (FAR)。 FAR 显示攻击模型将非训练样本(非成员)错误标记为训练(成员)样本的频率。高 FAR 使得 MI 攻击根本上不切实际,这对于成员推理等任务尤其重要,因为现实中大多数样本属于负(非训练)类。此外,我们表明,当前的 MI 攻击模型最多只能识别错误分类样本的成员资格,准确度也很一般,仅占训练样本的很小一部分。我们分析了一些以前没有全面探索过的成员资格推断新特征,包括到决策边界的距离和梯度范数,并得出结论,深度模型的响应在训练样本和非训练样本中大多相似。我们使用各种模型架构(包括 LeNet、AlexNet、ResNet 等)对图像分类任务(包括 MNIST、CIFAR-10、CIFAR-100 和 ImageNet)进行了多次实验。我们表明,即使攻击者获得了多种优势,当前最先进的 MI 攻击也无法同时实现高精度和低 FAR。源代码可在 https://github.com/shrezaei/MI-Attack 获取。
Recent studies propose membership inference (MI) attacks on deep models, where the goal is to infer if a sample has been used in the training process. Despite their apparent success, these studies only report accuracy, precision, and recall of the positive class (member class). Hence, the performance of these attacks have not been clearly reported on negative class (non-member class). In this paper, we show that the way the MI attack performance has been reported is often misleading because they suffer from high false positive rate or false alarm rate (FAR) that has not been reported. FAR shows how often the attack model mislabel non-training samples (non-member) as training (member) ones. The high FAR makes MI attacks fundamentally impractical, which is particularly more significant for tasks such as membership inference where the majority of samples in reality belong to the negative (non-training) class. Moreover, we show that the current MI attack models can only identify the membership of misclassified samples with mediocre accuracy at best, which only constitute a very small portion of training samples.We analyze several new features that have not been comprehensively explored for membership inference before, including distance to the decision boundary and gradient norms, and conclude that deep models’ responses are mostly similar among train and non-train samples. We conduct several experiments on image classification tasks, including MNIST, CIFAR-10, CIFAR-100, and ImageNet, using various model architecture, including LeNet, AlexNet, ResNet, etc. We show that the current state-of-the-art MI attacks cannot achieve high accuracy and low FAR at the same time, even when the attacker is given several advantages. The source code is available at https://github.com/shrezaei/MI-Attack.