Evaluating Robustness of Sequence-based Deepfake Detector Models by Adversarial Perturbation

Evaluating Robustness of Sequence-based Deepfake Detector Models by Adversarial Perturbation
复制标题

DOI:
10.1145/3494109.3527194
复制
发表时间:
2022-05
期刊:
Proceedings of the 1st Workshop on Security Implications of Deepfakes and Cheapfakes
影响因子:
--
通讯作者:
S. A. Shahriyar;M. Wright
S. A. Shahriyar;M. Wright
中科院分区:
其他
文献类型:
--
作者:
S. A. Shahriyar;M. Wright

文献摘要

相似文献

深度伪造视频的质量越来越好,并且可被用于危险的虚假信息传播活动。检测这些视频的迫切需求促使研究人员开发不同类型的检测模型。其中,利用时间信息的模型(即基于序列的模型)在检测方面比那些仅检测帧内差异的模型更有效。近期的研究表明,后者这种检测模型可能会被对抗样本愚弄,这利用了关于制作对抗性(静态)图像的丰富文献。然而,不太清楚的是这些攻击对基于多帧信息进行操作的基于序列的模型的效果如何。在本文中,我们探讨了快速梯度符号法(FGSM)和卡林尼 - 瓦格纳L2范数攻击在白盒和黑盒设置下愚弄基于序列的深度伪造检测模型的有效性。实验结果表明,这些攻击是有效的,在白盒和黑盒攻击场景中的最大成功率分别为99.72%和67.14%。这凸显了开发更强大的基于序列的深度伪造检测器的重要性,并为未来的研究开辟了方向。
Deepfake videos are getting better in quality and can be used for dangerous disinformation campaigns. The pressing need to detect these videos has motivated researchers to develop different types of detection models. Among them, the models that utilize temporal information (i.e., sequence-based models) are more effective at detection than the ones that only detect intra-frame discrepancies. Recent work has shown that the latter detection models can be fooled with adversarial examples, leveraging the rich literature on crafting adversarial (still) images. It is less clear, however, how well these attacks will work on sequence-based models that operate on information taken over multiple frames. In this paper, we explore the effectiveness of the Fast Gradient Sign Method (FGSM) and the Carlini-Wagner L2-norm attack to fool sequence-based deepfake detector models in both the white-box and black-box settings. The experimental results show that the attacks are effective with a maximum success rate of 99.72% and 67.14% in the white-box and black-box attack scenarios, respectively. This highlights the importance of developing more robust sequence-based deepfake detectors and opens up directions for future research.