Evaluating Robustness of Sequence-based Deepfake Detector Models by Adversarial Perturbation
Evaluating Robustness of Sequence-based Deepfake Detector Models by Adversarial Perturbation
复制标题
DOI:
10.1145/3494109.3527194
复制
发表时间:
2022-05
期刊:
影响因子:
--
通讯作者:
S. A. Shahriyar;M. Wright
中科院分区:
文献类型:
--
作者:
S. A. Shahriyar;M. Wright
Deepfake videos are getting better in quality and can be used for dangerous disinformation campaigns. The pressing need to detect these videos has motivated researchers to develop different types of detection models. Among them, the models that utilize temporal information (i.e., sequence-based models) are more effective at detection than the ones that only detect intra-frame discrepancies. Recent work has shown that the latter detection models can be fooled with adversarial examples, leveraging the rich literature on crafting adversarial (still) images. It is less clear, however, how well these attacks will work on sequence-based models that operate on information taken over multiple frames. In this paper, we explore the effectiveness of the Fast Gradient Sign Method (FGSM) and the Carlini-Wagner L2-norm attack to fool sequence-based deepfake detector models in both the white-box and black-box settings. The experimental results show that the attacks are effective with a maximum success rate of 99.72% and 67.14% in the white-box and black-box attack scenarios, respectively. This highlights the importance of developing more robust sequence-based deepfake detectors and opens up directions for future research.