Evading Watermark based Detection of AI-Generated Content

Evading Watermark based Detection of AI-Generated Content
复制标题

DOI:
10.1145/3576915.3623189
复制
发表时间:
2023-05
期刊:
Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Zhengyuan Jiang;Jinghuai Zhang;N. Gong
Zhengyuan Jiang;Jinghuai Zhang;N. Gong
中科院分区:
其他
文献类型:
--
作者:
Zhengyuan Jiang;Jinghuai Zhang;N. Gong

文献摘要

被引文献

相似文献

生成式人工智能模型可以生成极其逼真的内容,对信息的真实性构成越来越大的挑战。为了解决这些挑战,水印被用来检测AI生成的内容。具体而言,水印在发布之前嵌入到AI生成的内容中。一个内容被检测为AI生成的,如果一个类似的水印可以从它解码。在这项工作中,我们进行了系统的研究,这种基于水印的AI生成的内容检测的鲁棒性。我们专注于AI生成的图像。我们的工作表明,攻击者可以通过添加一个小的,人类无法感知的扰动,使后处理的图像逃避检测,同时保持其视觉质量的水印图像后处理。我们从理论上和经验上证明了我们的攻击的有效性。此外,为了逃避检测,我们的对抗性后处理方法向AI生成的图像添加了更小的扰动,从而比现有的流行后处理方法(如JPEG压缩,高斯模糊和亮度/对比度)更好地保持其视觉质量。我们的工作显示了现有基于水印的AI生成内容检测的不足,突出了新方法的迫切需求。我们的代码是公开的:https://github.com/zhengyuan-jiang/WEvade。
A generative AI model can generate extremely realistic-looking content, posing growing challenges to the authenticity of information. To address the challenges, watermark has been leveraged to detect AI-generated content. Specifically, a watermark is embedded into an AI-generated content before it is released. A content is detected as AI-generated if a similar watermark can be decoded from it. In this work, we perform a systematic study on the robustness of such watermark-based AI-generated content detection. We focus on AI-generated images. Our work shows that an attacker can post-process a watermarked image via adding a small, human-imperceptible perturbation to it, such that the post-processed image evades detection while maintaining its visual quality. We show the effectiveness of our attack both theoretically and empirically. Moreover, to evade detection, our adversarial post-processing method adds much smaller perturbations to AI-generated images and thus better maintain their visual quality than existing popular post-processing methods such as JPEG compression, Gaussian blur, and Brightness/Contrast. Our work shows the insufficiency of existing watermark-based detection of AI-generated content, highlighting the urgent needs of new methods. Our code is publicly available: https://github.com/zhengyuan-jiang/WEvade.