GenGradAttack: Efficient and Robust Targeted Adversarial Attacks Using Genetic Algorithms and Gradient-Based Fine-Tuning

GenGradAttack: Efficient and Robust Targeted Adversarial Attacks Using Genetic Algorithms and Gradient-Based Fine-Tuning
复制标题

DOI:
10.5220/0012314700003636
复制
发表时间:
2024
期刊:
--
影响因子:
--
通讯作者:
Naman Agarwal;James Pope
Naman Agarwal;James Pope
中科院分区:
其他
文献类型:
--
作者:
Naman Agarwal;James Pope

文献摘要

相似文献

对抗性攻击对机器学习模型的可靠性构成了严重威胁,可能会破坏实际应用中的信任。随着机器学习模型在自动驾驶汽车、医疗保健和金融等重要领域的部署,它们变得容易受到对抗性示例输入的影响,这些输入会导致错误的高置信度预测。这些攻击主要分为两类:白盒攻击,完全了解模型架构;黑盒攻击,对内部细节的访问有限或无法访问。本文介绍了一种在黑盒场景中进行针对性对抗攻击的新方法。通过结合遗传算法和基于梯度的微调,我们的方法有效地探索输入空间的扰动,而不需要访问内部模型的细节。随后,基于梯度的微调优化这些扰动,使它们与目标模型的决策边界对齐。这种双重策略旨在发展有效误导目标模型的扰动,同时最大限度地减少查询,确保隐形攻击。结果证明了GenGradAttack的有效性,在MNIST上实现了95.06%的对抗成功率(ASR),中位数查询计数为556。相比之下,传统的GenAttack实现了100%的ASR,但需要更多的查询。当应用于ImageNet上的InceptionV 3和Ens 4AdvInceptionV 3时,GenGradAttack的ASR分别为100%和96%,并且中位数查询更少。这些结果突出了我们的方法在生成具有减少查询计数的对抗性示例方面的效率和有效性,促进了我们对实际环境中对抗性漏洞的理解。
: Adversarial attacks pose a critical threat to the reliability of machine learning models, potentially undermining trust in practical applications. As machine learning models find deployment in vital domains like autonomous vehicles, healthcare, and finance, they become susceptible to adversarial examples—crafted inputs that induce erroneous high-confidence predictions. These attacks fall into two main categories: white-box, with full knowledge of model architecture, and black-box, with limited or no access to internal details. This paper introduces a novel approach for targeted adversarial attacks in black-box scenarios. By combining ge-netic algorithms and gradient-based fine-tuning, our method efficiently explores input space for perturbations without requiring access to internal model details. Subsequently, gradient-based fine-tuning optimizes these perturbations, aligning them with the target model’s decision boundary. This dual strategy aims to evolve perturbations that effectively mislead target models while minimizing queries, ensuring stealthy attacks. Results demonstrate the efficacy of GenGradAttack , achieving a remarkable 95.06% Adversarial Success Rate (ASR) on MNIST with a median query count of 556 . In contrast, conventional GenAttack achieved 100% ASR but required significantly more queries. When applied to InceptionV3 and Ens4AdvInceptionV3 on ImageNet, GenGradAttack outperformed GenAttack with 100% and 96% ASR, respectively, and fewer median queries. These results highlight the efficiency and effectiveness of our approach in generating adversarial examples with reduced query counts, advancing our understanding of adversarial vulnerabilities in practical contexts.