Character As Pixels: A Controllable Prompt Adversarial Attacking Framework for Black-Box Text Guided Image Generation Models

Character As Pixels: A Controllable Prompt Adversarial Attacking Framework for Black-Box Text Guided Image Generation Models
复制标题

DOI:
10.24963/ijcai.2023/109
复制
发表时间:
2023-08
期刊:
--
影响因子:
--
通讯作者:
Ziyi Kou;Shichao Pei;Yijun Tian;Xiangliang Zhang
Ziyi Kou;Shichao Pei;Yijun Tian;Xiangliang Zhang
中科院分区:
其他
文献类型:
--
作者:
Ziyi Kou;Shichao Pei;Yijun Tian;Xiangliang Zhang

文献摘要

被引文献

相似文献

在本文中,我们研究了黑盒场景中文本引导图像生成(Text2Image)模型的可控提示对抗攻击问题,其中目标是攻击特定的视觉主体(例如,将棕色狗改变为白色),通过轻微地(如果不是不可察觉的话)扰动驱动提示的字符(例如,"brown“到" br0wn”)。我们的研究是出于当前Text2Image攻击方法的局限性,这些方法仍然依赖于手动试验来创建对抗性提示。为了解决这些限制,我们开发CharGrad,字符级梯度为基础的攻击框架,取代特定字符的提示与像素级类似的交互式学习扰动方向的提示和更新的攻击考官生成的图像的基础上,一个新的代理扰动表示字符。我们使用两个公共图像字幕数据集的文本来评估CharGrad。实验结果表明,CharGrad在黑盒Text2Image模型下对生成图像的各种主题进行攻击时,性能优于现有的文本对抗攻击方法,且对提示符的扰动更小,攻击效率更高.
In this paper, we study a controllable prompt adversarial attacking problem for text guided image generation (Text2Image) models in the black-box scenario, where the goal is to attack specific visual subjects (e.g., changing a brown dog to white) in a generated image by slightly, if not imperceptibly, perturbing the characters of the driven prompt (e.g., ``brown'' to ``br0wn''). Our study is motivated by the limitations of current Text2Image attacking approaches that still rely on manual trials to create adversarial prompts. To address such limitations, we develop CharGrad, a character-level gradient based attacking framework that replaces specific characters of a prompt with pixel-level similar ones by interactively learning the perturbation direction for the prompt and updating the attacking examiner for the generated image based on a novel proxy perturbation representation for characters. We evaluate CharGrad using the texts from two public image captioning datasets. Results demonstrate that CharGrad outperforms existing text adversarial attacking approaches on attacking various subjects of generated images by black-box Text2Image models in a more effective and efficient way with less perturbation on the characters of the prompts.