Deep Partition Aggregation: Provable Defense against General Poisoning Attacks

Deep Partition Aggregation: Provable Defense against General Poisoning Attacks
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Alexander Levine;S. Feizi
Alexander Levine;S. Feizi
中科院分区:
其他
文献类型:
--
作者:
Alexander Levine;S. Feizi

文献摘要

被引文献

相似文献

对抗中毒攻击扭曲训练数据,以破坏分类器的测试时间行为。可证明的防御提供了每个测试样本的证书,这是训练集的任何对抗性扭曲的大小,可能会破坏测试样本的分类。我们提出了两种可证明的防御中毒攻击的防御措施:(i)对一般中毒威胁模型的认证辩护深度分区聚合(DPA),定义为插入或删除有限数量的样本到培训集中 - 暗示,此威胁模型还包括对有限数量的图像和/或标签的任意扭曲; (ii)半监督的DPA(SS-DPA),这是针对贴标签的中毒攻击的认证辩护。 DPA是一种合奏方法,在该方法中,对基本模型进行了对由哈希功能确定的训练集的分区进行训练。 DPA与子集汇总有关,这是一种经典机器学习中的经过良好研究的合奏方法。 DPA也可以看作是随机消融(Levine&Feizi,2020a)的扩展,这是对稀疏逃避攻击的认证辩护,以延伸到中毒领域。我们的标签上夹板防御SS-DPA使用半监督的学习算法作为其基本分类器模型:我们使用整个未标记的培训集训练每个基本分类器,此外还包括分区标签。 SS-DPA的表现优于现有的贴标签攻击的认证防御(Rosenfeld等,2020)。 SS-DPA认证> = = 50%的测试图像相对于675个标签翻转(Vs. = 500> 500毒图像插入MNIST上的测试图像的50%,以及在CIFAR-10上插入的9个。这些结果确定了新的最新最先进 - 可证明防御毒物攻击的防御能力。
Adversarial poisoning attacks distort training data in order to corrupt the test-time behavior of a classifier. A provable defense provides a certificate for each test sample, which is a lower bound on the magnitude of any adversarial distortion of the training set that can corrupt the test sample's classification. We propose two provable defenses against poisoning attacks: (i) Deep Partition Aggregation (DPA), a certified defense against a general poisoning threat model, defined as the insertion or deletion of a bounded number of samples to the training set -- by implication, this threat model also includes arbitrary distortions to a bounded number of images and/or labels; and (ii) Semi-Supervised DPA (SS-DPA), a certified defense against label-flipping poisoning attacks. DPA is an ensemble method where base models are trained on partitions of the training set determined by a hash function. DPA is related to subset aggregation, a well-studied ensemble method in classical machine learning. DPA can also be viewed as an extension of randomized ablation (Levine & Feizi, 2020a), a certified defense against sparse evasion attacks, to the poisoning domain. Our label-flipping defense, SS-DPA, uses a semi-supervised learning algorithm as its base classifier model: we train each base classifier using the entire unlabeled training set in addition to the labels for a partition. SS-DPA outperforms the existing certified defense for label-flipping attacks (Rosenfeld et al., 2020). SS-DPA certifies >= 50% of test images against 675 label flips (vs. = 50% of test images against > 500 poison image insertions on MNIST, and nine insertions on CIFAR-10. These results establish new state-of-the-art provable defenses against poison attacks.