Speech Privacy Leakage from Shared Gradients in Distributed Learning

Speech Privacy Leakage from Shared Gradients in Distributed Learning
复制标题

DOI:
10.1109/icassp49357.2023.10095443
复制
发表时间:
2023-02
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Zhuohang Li;Jiaxin Zhang;Jian Liu
Zhuohang Li;Jiaxin Zhang;Jian Liu
中科院分区:
其他
文献类型:
--
作者:
Zhuohang Li;Jiaxin Zhang;Jian Liu

文献摘要

被引文献

相似文献

分布式机器学习范例,如联合学习,最近已被用于语音分析的许多隐私关键型应用中。然而,这样的框架很容易受到来自共享梯度的隐私泄露攻击。尽管在图像领域进行了广泛的研究,但从梯度角度对语音隐私泄漏的探索还是相当有限的。在本文中,我们探索了在分布式学习环境中从共享梯度中恢复私人语音/说话人信息的方法。我们在具有两种不同语音特征的关键词检测模型上进行了实验,通过测量原始语音信号和恢复语音信号之间的相似度来量化泄漏的信息量。我们进一步论证了在不访问用户数据的情况下,在分布式学习框架下推断包括语音内容和说话人身份在内的各种级别的旁路信息的可行性。
Distributed machine learning paradigms, such as federated learning, have been recently adopted in many privacy-critical applications for speech analysis. However, such frameworks are vulnerable to privacy leakage attacks from shared gradients. Despite extensive efforts in the image domain, the exploration of speech privacy leakage from gradients is quite limited. In this paper, we explore methods for recovering private speech/speaker information from the shared gradients in distributed learning settings. We conduct experiments on a keyword spotting model with two different types of speech features to quantify the amount of leaked information by measuring the similarity between the original and recovered speech signals. We further demonstrate the feasibility of inferring various levels of side-channel information, including speech content and speaker identity, under the distributed learning framework without accessing the user’s data.