VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive Perturbation

VSMask: Defending Against Voice Synthesis Attack via Real-Time Predictive Perturbation
复制标题

DOI:
10.1145/3558482.3590189
复制
发表时间:
2023-05
期刊:
Proceedings of the 16th ACM Conference on Security and Privacy in Wireless and Mobile Networks
影响因子:
--
通讯作者:
Yuanda Wang;Hanqing Guo;Guangjing Wang;Bocheng Chen;Qiben Yan
Yuanda Wang;Hanqing Guo;Guangjing Wang;Bocheng Chen;Qiben Yan
中科院分区:
其他
文献类型:
--
作者:
Yuanda Wang;Hanqing Guo;Guangjing Wang;Bocheng Chen;Qiben Yan

文献摘要

被引文献

相似文献

基于深度学习的语音合成技术生成的人工语音已被用于深度假冒或身份盗窃攻击。现有的防御机制在原始语音音频中注入微妙的对抗性扰动,以误导语音合成模型。然而,优化对抗性扰动不仅需要消耗大量的计算时间,而且还需要整个语音的可用性。因此,它们不适合保护实时语音流,例如语音消息或在线会议。本文提出了一种针对语音合成攻击的实时防护机制VSMASK。与离线保护方案不同,VSMAsk利用预测神经网络来预测即将到来的流传输语音的最有效扰动。VSMask引入了一种为任意语音输入量身定做的通用扰动,以完整地屏蔽实时语音。为了最大限度地减少受保护语音中的音频失真,我们实现了基于权重的扰动约束来降低附加扰动的可感知性。我们综合评估了不同场景下的VSMASK保护性能。实验结果表明,VSMASK能够有效防御3种流行的语音合成模型。任何合成语音都无法欺骗说话人验证模型或具有VSMASK保护的人耳。在物理世界的实验中,我们演示了VSMAsk通过在空中注入扰动来成功地保护实时语音。
Deep learning based voice synthesis technology generates artificial human-like speeches, which has been used in deepfakes or identity theft attacks. Existing defense mechanisms inject subtle adversarial perturbations into the raw speech audios to mislead the voice synthesis models. However, optimizing the adversarial perturbation not only consumes substantial computation time, but it also requires the availability of entire speech. Therefore, they are not suitable for protecting live speech streams, such as voice messages or online meetings. In this paper, we propose VSMask, a real-time protection mechanism against voice synthesis attacks. Different from offline protection schemes, VSMask leverages a predictive neural network to forecast the most effective perturbation for the upcoming streaming speech. VSMask introduces a universal perturbation tailored for arbitrary speech input to shield a real-time speech in its entirety. To minimize the audio distortion within the protected speech, we implement a weight-based perturbation constraint to reduce the perceptibility of the added perturbation. We comprehensively evaluate VSMask protection performance under different scenarios. The experimental results indicate that VSMask can effectively defend against 3 popular voice synthesis models. None of the synthetic voice could deceive the speaker verification models or human ears with VSMask protection. In a physical world experiment, we demonstrate that VSMask successfully safeguards the real-time speech by injecting the perturbation over the air.