VQVC+: One-Shot Voice Conversion by Vector Quantization and U-Net architecture
VQVC+: One-Shot Voice Conversion by Vector Quantization and U-Net architecture
复制标题
VQVC:通过矢量量化和 U-Net 架构进行一次性语音转换
DOI:
10.21437/interspeech.2020-1443
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Hung
中科院分区:
文献类型:
--
作者:
Da;Yen;Hung
Voice conversion (VC) is a task that transforms the source speaker's timbre, accent, and tones in audio into another one's while preserving the linguistic content. It is still a challenging work, especially in a one-shot setting. Auto-encoder-based VC methods disentangle the speaker and the content in input speech without given the speaker's identity, so these methods can further generalize to unseen speakers. The disentangle capability is achieved by vector quantization (VQ), adversarial training, or instance normalization (IN). However, the imperfect disentanglement may harm the quality of output speech. In this work, to further improve audio quality, we use the U-Net architecture within an auto-encoder-based VC system. We find that to leverage the U-Net architecture, a strong information bottleneck is necessary. The VQ-based method, which quantizes the latent vectors, can serve the purpose. The objective and the subjective evaluations show that the proposed method performs well in both audio naturalness and speaker similarity.
DOI:
10.1109/icassp40776.2020.9054734
发表时间:
2020-04
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
Kaizhi Qian;Zeyu Jin;M. Hasegawa-Johnson;G. J. Mysore
通讯作者:
Kaizhi Qian;Zeyu Jin;M. Hasegawa-Johnson;G. J. Mysore