Privacy and Utility of X-Vector Based Speaker Anonymization

Privacy and Utility of X-Vector Based Speaker Anonymization
复制标题

DOI:
10.1109/taslp.2022.3190741
复制
发表时间:
2022
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
B. M. L. Srivastava;Mohamed Maouche;Md. Sahidullah;E. Vincent;A. Bellet;M. Tommasi;N. Tomashenko;Xin Wang;J. Yamagishi
B. M. L. Srivastava;Mohamed Maouche;Md. Sahidullah;E. Vincent;A. Bellet;M. Tommasi;N. Tomashenko;Xin Wang;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
B. M. L. Srivastava;Mohamed Maouche;Md. Sahidullah;E. Vincent;A. Bellet;M. Tommasi;N. Tomashenko;Xin Wang;J. Yamagishi

文献摘要

被引文献

相似文献

我们研究的情况下,个人(扬声器)有助于出版的匿名语音语料库。数据用户利用这个公共语料库进行下游任务,例如,训练自动语音识别(ASR)系统,而攻击者可能会尝试使用辅助知识对其进行去匿名化。在这种情况下,说话人匿名的目的是隐藏说话人身份,同时保持语音数据的质量和有用性。在这篇文章中,我们研究了基于x向量的说话人匿名化,这是语音隐私挑战赛中的领先方法,它将说话人的声音转换为随机伪说话人的声音。我们发现,匿名化的强度显着不同,这取决于如何选择伪扬声器。我们探索了这一步的四个设计选择:扬声器之间的距离度量,扬声器空间的区域,伪扬声器被选中,其性别,以及是否将其分配给原始扬声器的一个或所有话语。我们从我们的威胁模型中涉及的三个参与者的角度评估匿名化的质量,即扬声器,用户和攻击者。为了衡量隐私和效用,我们分别使用攻击者获得的可链接性得分和在匿名数据上训练的ASR模型获得的解码单词错误率。LibriSpeech上的实验表明,设计选择的最佳组合在隐私和实用性方面都具有最先进的性能。在Mozilla Common Voice上的实验进一步表明,它保证了50个说话人之间的匿名化水平与20,000个说话人之间的原始语音相同。
We study the scenario where individuals (speakers) contribute to the publication of an anonymized speech corpus. Data users leverage this public corpus for downstream tasks, e.g., training an automatic speech recognition (ASR) system, while attackers may attempt to de-anonymize it using auxiliary knowledge. Motivated by this scenario, speaker anonymization aims to conceal speaker identity while preserving the quality and usefulness of speech data. In this article, we study x-vector based speaker anonymization, the leading approach in the VoicePrivacy Challenge, which converts the speaker's voice into that of a random pseudo-speaker. We show that the strength of anonymization varies significantly depending on how the pseudo-speaker is chosen. We explore four design choices for this step: the distance metric between speakers, the region of speaker space where the pseudo-speaker is picked, its gender, and whether to assign it to one or all utterances of the original speaker. We assess the quality of anonymization from the perspective of the three actors involved in our threat model, namely the speaker, the user and the attacker. To measure privacy and utility, we use respectively the linkability score achieved by the attackers and the decoding word error rate achieved by an ASR model trained on the anonymized data. Experiments on LibriSpeech show that the best combination of design choices yields state-of-the-art performance in terms of both privacy and utility. Experiments on Mozilla Common Voice further show that it guarantees the same anonymization level against re-identification attacks among 50 speakers as original speech among 20,000 speakers.