Hiding Speaker’s Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline

Hiding Speaker’s Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline
复制标题

DOI:
10.1109/icassp49357.2023.10096749
复制
发表时间:
2022-11
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Paul-Gauthier Noé;Xiaoxiao Miao;Xin Wang;J. Yamagishi;J. Bonastre;D. Matrouf
Paul-Gauthier Noé;Xiaoxiao Miao;Xin Wang;J. Yamagishi;J. Bonastre;D. Matrouf
中科院分区:
其他
文献类型:
--
作者:
Paul-Gauthier Noé;Xiaoxiao Miao;Xin Wang;J. Yamagishi;J. Bonastre;D. Matrouf

文献摘要

相似文献

在分析/合成管道中使用现代声码器使我们能够研究可用于隐私目的的高质量语音转换。在这里,我们建议变换扬声器嵌入和音高,以隐藏扬声器的性别。基于ECAPA-TDNN的说话人表示馈入HiFiGAN声码器使用神经判别分析方法进行保护,这与隐私的零证据概念一致。这种方法大大减少了语音中与说话者性别有关的信息,同时保留了语音内容和由此产生的受保护声音的一致性。
The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representation fed into a HiFiGAN vocoder is protected using a neural-discriminant analysis approach, which is consistent with the zero-evidence concept of privacy. This approach significantly reduces the information in speech related to the speaker’s sex while preserving speech content and some consistency in the resulting protected voices.