Learning and Certification under Instance-targeted Poisoning

Learning and Certification under Instance-targeted Poisoning
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Ji Gao;Amin Karbasi;Mohammad Mahmoody
Ji Gao;Amin Karbasi;Mohammad Mahmoody
中科院分区:
其他
文献类型:
--
作者:
Ji Gao;Amin Karbasi;Mohammad Mahmoody

文献摘要

相似文献

在本文中,我们研究了PAC的可学习性和在实例定为实例的中毒攻击下进行预测的认证,在这种情况下,知道测试实例的对手可能会改变培训设置的一部分,目的是在测试实例中欺骗学习者。我们的第一个贡献是在各种环境中形式化问题,并明确模拟微妙的方面,例如学习的适当或不当性质,学习者的随机性以及(是否)对手的攻击是否可以取决于它。我们的主要结果表明,当对手的预算随着样本复杂性而倍增时,(不当)PAC可学习性和认证就可以实现;相反,当对手的预算随样品复杂性线性增长时,对手可能会使预期的0-1损失提高到一个。我们还以相同的攻击模型研究了特定于分布的PAC学习,并表明在自然分布下学习一半空间是可能的。最后,我们在实证上研究了K最近的邻居,逻辑回归,多层感知器和卷积神经网络的鲁棒性,以针对有针对性的伪造攻击。我们的实验结果表明,许多模型,尤其是最新的神经网络,确实容易受到这些强烈攻击的影响。有趣的是,我们观察到具有高标准精度的方法可能更容易受到实例定位的中毒攻击的影响。
In this paper, we study PAC learnability and certification of predictions under instance-targeted poisoning attacks, where the adversary who knows the test instance may change a fraction of the training set with the goal of fooling the learner at the test instance. Our first contribution is to formalize the problem in various settings and to explicitly model subtle aspects such as the proper or improper nature of the learning, learner's randomness, and whether (or not) adversary's attack can depend on it. Our main result shows that when the budget of the adversary scales sublinearly with the sample complexity, (improper) PAC learnability and certification are achievable; in contrast, when the adversary's budget grows linearly with the sample complexity, the adversary can potentially drive up the expected 0-1 loss to one. We also study distribution-specific PAC learning in the same attack model and show that proper learning with certification is possible for learning half spaces under natural distributions. Finally, we empirically study the robustness of K nearest neighbour, logistic regression, multi-layer perceptron, and convolutional neural network on real data sets against targeted-poisoning attacks. Our experimental results show that many models, especially state-of-the-art neural networks, are indeed vulnerable to these strong attacks. Interestingly, we observe that methods with high standard accuracy might be more vulnerable to instance-targeted poisoning attacks.