Speaker anonymization by pitch shifting based on time-scale modification

Speaker anonymization by pitch shifting based on time-scale modification
复制标题

DOI:
10.21437/spsc.2022-7
复制
发表时间:
2022-09
期刊:
2nd Symposium on Security and Privacy in Speech Communication
影响因子:
--
通讯作者:
Candy Olivia Mawalim;S. Okada;M. Unoki
Candy Olivia Mawalim;S. Okada;M. Unoki
中科院分区:
其他
文献类型:
--
作者:
Candy Olivia Mawalim;S. Okada;M. Unoki

文献摘要

相似文献

数字技术中语音的使用越来越多,这引起了隐私问题,因为语音包含生物特征信息。已经提出了几种处理这个问题的方法,包括说话人匿名化或去识别。说话人匿名化的目的是在保留语音的其他属性(包括语言信息)的同时,抑制个人身份信息(PII)。在这项研究中,我们利用时间尺度修改(TSM)语音信号处理的说话人匿名。语音信号处理方法比最先进的基于x向量的说话人匿名化方法复杂得多,因为它不需要训练过程。我们提出了匿名化方法,使用两个主要类别的TSM,同步的语音添加(SOLA)为基础的算法和相位声码器为基础的TSM(PV-TSM)。为了评估我们提出的方法,我们利用标准的客观评估中引入的语音隐私的挑战。结果表明,我们基于PV-TSM的方法比基线系统更好地平衡了隐私和实用性指标,特别是在匿名注册和匿名试验(a-a)中使用自动说话人验证(ASV)系统进行评估时。此外,我们的方法优于基于x向量的说话人方法,后者在训练过程复杂、a-a场景中隐私性低以及语音独特性低方面存在局限性。
The increasing usage of speech in digital technology raises a privacy issue because speech contains biometric information. Several methods of dealing with this issue have been proposed, including speaker anonymization or de-identification. Speaker anonymization aims to suppress personally identifiable information (PII) while keeping the other speech properties, including linguistic information. In this study, we utilize time-scale modification (TSM) speech signal processing for speaker anonymization. Speech signal processing approaches are significantly less complex than the state-of-the-art x-vector-based speaker anonymization method because it does not require a training process. We propose anonymization methods using two major categories of TSM, synchronous overlap-add (SOLA)- based algorithm and phase vocoder-based TSM (PV-TSM). For evaluating our proposed methods, we utilize the standard objec-tive evaluation introduced in the VoicePrivacy challenge. The results show that our method based on the PV-TSM balances privacy and utility metrics better than baseline systems, especially when evaluating with an automatic speaker verification (ASV) system in anonymized enrollment and anonymized trials (a-a). Further, our method outperformed the x-vector-based speaker method, which has limitations in its complex training process, low privacy in an a-a scenario, and low voice distinctiveness.