Deepfake Video Detection Based on Spatial, Spectral, and Temporal Inconsistencies Using Multimodal Deep Learning

Deepfake Video Detection Based on Spatial, Spectral, and Temporal Inconsistencies Using Multimodal Deep Learning
复制标题

DOI:
10.1109/aipr50011.2020.9425167
复制
发表时间:
2020-10
期刊:
2020 IEEE Applied Imagery Pattern Recognition Workshop (AIPR)
影响因子:
--
通讯作者:
John Lewis;Imad Eddine Toubal;Helen Chen;Vishal Sandesera;M. Lomnitz;Z. Hampel-Arias;P. Calyam;K. Palaniappan
John Lewis;Imad Eddine Toubal;Helen Chen;Vishal Sandesera;M. Lomnitz;Z. Hampel-Arias;P. Calyam;K. Palaniappan
中科院分区:
其他
文献类型:
--
作者:
John Lewis;Imad Eddine Toubal;Helen Chen;Vishal Sandesera;M. Lomnitz;Z. Hampel-Arias;P. Calyam;K. Palaniappan

文献摘要

被引文献

相似文献

数字媒体的认证已经成为现代社会日益迫切的需要。自生成对抗网络(GAN)引入以来,合成媒体变得越来越难以识别。包含改变的人脸和/或声音的合成视频被称为deepfakes,威胁到数字媒体的信任和隐私。深度造假可以被武器化,用于政治利益、诽谤和破坏公众人物的声誉。尽管deepfake存在缺陷,但人们很难区分真实的和被操纵的图像和视频。因此,重要的是要有自动化系统,准确和有效地分类数字内容的有效性。许多最近的deepfake检测方法使用单帧视频,并专注于图像中的空间信息来推断视频的真实性。一些有前途的方法利用操纵视频的时间不一致性,然而,研究主要集中在空间特征。我们提出了一种混合深度学习方法,该方法使用以一致方式耦合的空间,光谱和时间内容来区分真实的和假视频。我们表明,离散余弦变换可以通过捕获单个帧的频谱特征来提高深度伪造检测。在这项工作中,我们构建了一个多模态网络,探索了检测deepfake视频的新功能,在Facebook Deepfake检测挑战(DFDC)数据集上实现了61.95%的准确率。
Authentication of digital media has become an ever-pressing necessity for modern society. Since the introduction of Generative Adversarial Networks (GANs), synthetic media has become increasingly difficult to identify. Synthetic videos that contain altered faces and/or voices of a person are known as deepfakes and threaten trust and privacy in digital media. Deep-fakes can be weaponized for political advantage, slander, and to undermine the reputation of public figures. Despite imperfections of deepfakes, people struggle to distinguish between authentic and manipulated images and videos. Consequently, it is important to have automated systems that accurately and efficiently classify the validity of digital content. Many recent deepfake detection methods use single frames of video and focus on the spatial information in the image to infer the authenticity of the video. Some promising approaches exploit the temporal inconsistencies of manipulated videos; however, research primarily focuses on spatial features. We propose a hybrid deep learning approach that uses spatial, spectral, and temporal content that is coupled in a consistent way to differentiate real and fake videos. We show that the Discrete Cosine transform can improve deepfake detection by capturing spectral features of individual frames. In this work, we build a multimodal network that explores new features to detect deepfake videos, achieving 61.95% accuracy on the Facebook Deepfake Detection Challenge (DFDC) dataset.