F0-Consistent Many-To-Many Non-Parallel Voice Conversion Via Conditional Autoencoder

F0-Consistent Many-To-Many Non-Parallel Voice Conversion Via Conditional Autoencoder
复制标题

DOI:
10.1109/icassp40776.2020.9054734
复制
发表时间:
2020-04
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Kaizhi Qian;Zeyu Jin;M. Hasegawa-Johnson;G. J. Mysore
Kaizhi Qian;Zeyu Jin;M. Hasegawa-Johnson;G. J. Mysore
中科院分区:
其他
文献类型:
--
作者:
Kaizhi Qian;Zeyu Jin;M. Hasegawa-Johnson;G. J. Mysore

文献摘要

被引文献

相似文献

非并行多对多语音转换仍然是一项有趣但具有挑战性的语音处理任务。已经提出了许多基于风格迁移的方法,如生成性对抗网络(GANS)和变分自动编码器(VAEs)。最近,基于条件自动编码器(CAES)的AU-TOVC方法通过利用信息约束瓶颈来解开说话人身份和语音内容的纠缠,并通过替换不同说话人身份嵌入来合成新的语音来实现零镜头转换。然而,我们发现,在说话人身份从语音内容中解脱出来的同时,大量的韵律信息,如源F0,通过瓶颈泄漏,导致目标F0不自然地波动。此外,AutoVC无法控制转换后的F0,因此不适合许多应用。在本文中,我们对基于自动编码的语音转换进行了修改和改进,以同时分离内容、F0和说话人身份。因此,我们可以控制F0轮廓,生成与目标说话人一致的F0语音,显著提高质量和相似度。我们通过定量和定性分析来支持我们的改进。
Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Many style-transfer-inspired methods such as generative adversarial networks (GANs) and variational autoencoders (VAEs) have been proposed. Recently, AU-TOVC, a conditional autoencoders (CAEs) based method achieved state-of-the-art results by disentangling the speaker identity and speech content using information-constraining bottlenecks, and it achieves zero-shot conversion by swapping in a different speaker’s identity embedding to synthesize a new voice. However, we found that while speaker identity is disentangled from speech content, a significant amount of prosodic information, such as source F0, leaks through the bottleneck, causing target F0 to fluctuate unnaturally. Furthermore, AutoVC has no control of the converted F0 and thus unsuitable for many applications. In the paper, we modified and improved autoencoder-based voice conversion to disentangle content, F0, and speaker identity at the same time. Therefore, we can control the F0 contour, generate speech with F0 consistent with the target speaker, and significantly improve quality and similarity. We support our improvement through quantitative and qualitative analysis.