DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake Voices

DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake Voices
复制标题

DOI:
10.1145/3394171.3413716
复制
发表时间:
2020-05
期刊:
Proceedings of the 28th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Run Wang;Felix Juefei-Xu;Yihao Huang;Qing Guo;Xiaofei Xie;L. Ma;Yang Liu
Run Wang;Felix Juefei-Xu;Yihao Huang;Qing Guo;Xiaofei Xie;L. Ma;Yang Liu
中科院分区:
其他
文献类型:
--
作者:
Run Wang;Felix Juefei-Xu;Yihao Huang;Qing Guo;Xiaofei Xie;L. Ma;Yang Liu

文献摘要

被引文献

相似文献

随着语音合成的最新进展,人工智能合成的假语音对人耳来说是无法区分的,并被广泛应用于产生逼真和自然的DeepFake,对我们的社会表现出真实的威胁。然而,用于合成假声音的有效和强大的检测器仍处于起步阶段,尚未准备好完全解决这一新兴威胁。在本文中,我们设计了一种新的方法,命名为DeepSonar,基于监测说话人识别(SR)系统的神经元行为,即,深度神经网络(DNN),用于识别人工智能合成的假声音。逐层神经元行为提供了一个重要的洞察力,可以细致地捕捉输入之间的差异,这些差异被广泛用于构建安全,鲁棒和可解释的DNN。在这项工作中,我们利用逐层神经元激活模式的力量,推测它们可以捕捉真实的和人工智能合成的假声音之间的细微差异,为分类器提供比原始输入更清晰的信号。在三个包含英语和汉语的数据集(包括Google,百度等商业产品)上进行了实验,以证实DeepSonar在识别假语音方面的高检测率(98.1%的平均准确率)和低误报率(约2%的错误率)。此外,大量的实验结果也证明了它对操纵攻击的鲁棒性(例如,语音转换和附加的真实世界噪声)。我们的工作进一步提出了一种新的见解,即采用神经元行为进行有效和强大的人工智能辅助多媒体假货取证,作为一种由内而外的方法,而不是受到合成假货中引入的各种工件的激励和影响。
With the recent advances in voice synthesis, AI-synthesized fake voices are indistinguishable to human ears and widely are applied to produce realistic and natural DeepFakes, exhibiting real threats to our society. However, effective and robust detectors for synthesized fake voices are still in their infancy and are not ready to fully tackle this emerging threat. In this paper, we devise a novel approach, named DeepSonar, based on monitoring neuron behaviors of speaker recognition (SR) system, i.e., a deep neural network (DNN), to discern AI-synthesized fake voices. Layer-wise neuron behaviors provide an important insight to meticulously catch the differences among inputs, which are widely employed for building safety, robust, and interpretable DNNs. In this work, we leverage the power of layer-wise neuron activation patterns with a conjecture that they can capture the subtle differences between real and AI-synthesized fake voices, in providing a cleaner signal to classifiers than raw inputs. Experiments are conducted on three datasets (including commercial products from Google, Baidu, etc) containing both English and Chinese languages to corroborate the high detection rates (98.1% average accuracy) and low false alarm rates (about 2% error rate) of DeepSonar in discerning fake voices. Furthermore, extensive experimental results also demonstrate its robustness against manipulation attacks (e.g., voice conversion and additive real-world noises). Our work further poses a new insight into adopting neuron behaviors for effective and robust AI aided multimedia fakes forensics as an inside-out approach instead of being motivated and swayed by various artifacts introduced in synthesizing fakes.