Run-Off Election: Improved Provable Defense against Data Poisoning Attacks

Run-Off Election: Improved Provable Defense against Data Poisoning Attacks
复制标题

DOI:
10.48550/arxiv.2302.02300
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Keivan Rezaei;Kiarash Banihashem;A. Chegini;S. Feizi
Keivan Rezaei;Kiarash Banihashem;A. Chegini;S. Feizi
中科院分区:
其他
文献类型:
--
作者:
Keivan Rezaei;Kiarash Banihashem;A. Chegini;S. Feizi

文献摘要

相似文献

在数据中毒攻击中,攻击者试图通过添加、修改或删除训练数据中的样本来改变模型的预测。最近,已经提出了基于集成的方法来获得针对数据中毒的可证明防御,其中通过在多个基本模型中进行多数投票来进行预测。在这项工作中,我们表明,仅仅考虑集成防御中的多数投票是浪费的,因为它不能有效地利用基本模型的逻辑层中的可用信息。取而代之的是,我们提出了一种基于两轮基本模型选举的新型聚合方法——决胜选举(ROE):在第一轮中,模型投票给他们喜欢的类别,然后在第一轮中排名前两位的类别之间进行第二轮决胜选举。在此基础上,我们提出了基于深度分割聚合(DPA)和有限聚合(FA)方法的DPA+ROE和FA+ROE防御方法。我们在MNIST、CIFAR-10和GTSRB上评估了我们的方法,并获得了高达3%-4%的认证精度改进。此外,通过将ROE应用于增强版本的DPA,与当前最先进的技术相比,我们获得了约12%-27%的改进,建立了一个新的最先进的(点)认证的抗数据中毒的鲁棒性。在许多情况下,我们的方法优于最先进的技术,即使使用的计算能力只有原来的32倍。
In data poisoning attacks, an adversary tries to change a model's prediction by adding, modifying, or removing samples in the training data. Recently, ensemble-based approaches for obtaining provable defenses against data poisoning have been proposed where predictions are done by taking a majority vote across multiple base models. In this work, we show that merely considering the majority vote in ensemble defenses is wasteful as it does not effectively utilize available information in the logits layers of the base models. Instead, we propose Run-Off Election (ROE), a novel aggregation method based on a two-round election across the base models: In the first round, models vote for their preferred class and then a second, Run-Off election is held between the top two classes in the first round. Based on this approach, we propose DPA+ROE and FA+ROE defense methods based on Deep Partition Aggregation (DPA) and Finite Aggregation (FA) approaches from prior work. We evaluate our methods on MNIST, CIFAR-10, and GTSRB and obtain improvements in certified accuracy by up to 3%-4%. Also, by applying ROE on a boosted version of DPA, we gain improvements around 12%-27% comparing to the current state-of-the-art, establishing a new state-of-the-art in (pointwise) certified robustness against data poisoning. In many cases, our approach outperforms the state-of-the-art, even when using 32 times less computational power.