UniAP: Protecting Speech Privacy With Non-Targeted Universal Adversarial Perturbations

UniAP: Protecting Speech Privacy With Non-Targeted Universal Adversarial Perturbations
复制标题

DOI:
10.1109/tdsc.2023.3242292
复制
发表时间:
2024-01
影响因子:
7.3
通讯作者:
Peng Cheng;Yuexin Wu;Yuan Hong;Zhongjie Ba;Feng Lin;Liwang Lu;Kui Ren
Peng Cheng;Yuexin Wu;Yuan Hong;Zhongjie Ba;Feng Lin;Liwang Lu;Kui Ren
中科院分区:
计算机科学2区
文献类型:
--
作者:
Peng Cheng;Yuexin Wu;Yuan Hong;Zhongjie Ba;Feng Lin;Liwang Lu;Kui Ren

文献摘要

相似文献

智能设备上无处不在的麦克风极大地增加了用户对语音隐私的担忧。由于麦克风主要由硬件/软件开发商控制,利润驱动的组织可以通过深度学习模型轻松地大规模收集和分析个人的日常对话,用户无法阻止此类侵犯隐私的行为。在本文中,我们建议UNIAP使用户能够在不影响其日常语音活动的情况下保护其语音隐私不受大规模分析的影响。基于对识别模型的观察,我们利用对抗性学习产生准不可察觉的扰动来干扰附近麦克风捕获的语音信号,从而将录音的识别结果混淆为无意义的内容。实验证明,无论用户说什么、什么时候说,我们的扰动都可以保护用户隐私。通过训练优化进一步提高了干扰性能的稳定性。此外,该扰动对噪声去除技术具有很强的鲁棒性。广泛的评估表明,在现实聊天场景中,我们的扰动在数字域获得了超过87%的成功干扰率,在普通和具有挑战性的场景中分别达到了至少90%和70%的成功率。此外,我们的扰动,仅在DeepSpeech上训练,与基于类似架构的其他模型相比,显示出良好的可移植性。
Ubiquitous microphones on smart devices considerably raise users’ concerns about speech privacy. Since the microphones are primarily controlled by hardware/software developers, profit-driven organizations can easily collect and analyze individuals’ daily conversations on a large scale with deep learning models, and users have no means to stop such privacy-violating behavior. In this article, we propose UniAP to empower users with the capability of protecting their speech privacy from the large-scale analysis without affecting their routine voice activities. Based on our observation of the recognition model, we utilize adversarial learning to generate quasi-imperceptible perturbations to disturb speech signals captured by nearby microphones, thus obfuscating the recognition results of recordings into meaningless contents. As validated in experiments, our perturbations can protect user privacy regardless of what users speak and when they speak. The jamming performance stability is further improved by training optimization. Additionally, the perturbations are robust against noise removal techniques. Extensive evaluations show that our perturbations achieve successful jamming rates of more than 87% in the digital domain and at least 90% and 70% for common and challenging settings, respectively, in the real-life chatting scenario. Moreover, our perturbations, solely trained on DeepSpeech, exhibit good transferability over other models based on similar architecture.