VQVC+: One-Shot Voice Conversion by Vector Quantization and U-Net architecture

VQVC+: One-Shot Voice Conversion by Vector Quantization and U-Net architecture
复制标题

VQVC:通过矢量量化和 U-Net 架构进行一次性语音转换

DOI:
10.21437/interspeech.2020-1443
复制
发表时间:
2020
期刊:
ArXiv
影响因子:
--
通讯作者:
Hung
Hung
中科院分区:
--
文献类型:
--
作者:
Da;Yen;Hung

文献摘要

参考文献

被引文献

相似文献

语音转换是在保留语言内容的同时,将源说话人的音色、口音和音调转换成另一种音色、口音和音调的任务。这仍然是一项具有挑战性的工作,特别是在一次拍摄的背景下。基于自动编码的VC方法在不给出说话人身份的情况下,将说话人和输入语音中的内容分开,从而可以进一步推广到看不见的说话人。解缠能力通过矢量量化(VQ)、对抗性训练或实例归一化(IN)来实现。然而,不完善的解缠可能会损害输出语音的质量。在这项工作中,为了进一步提高音频质量,我们在一个基于自动编码器的VC系统中使用了U-Net架构。我们发现,要利用U-Net体系结构,强大的信息瓶颈是必要的。基于矢量量化的方法对潜在向量进行量化,可以达到这一目的。客观评价和主观评价表明,该方法在音频自然度和说话人相似度方面都有较好的表现。
Voice conversion (VC) is a task that transforms the source speaker's timbre, accent, and tones in audio into another one's while preserving the linguistic content. It is still a challenging work, especially in a one-shot setting. Auto-encoder-based VC methods disentangle the speaker and the content in input speech without given the speaker's identity, so these methods can further generalize to unseen speakers. The disentangle capability is achieved by vector quantization (VQ), adversarial training, or instance normalization (IN). However, the imperfect disentanglement may harm the quality of output speech. In this work, to further improve audio quality, we use the U-Net architecture within an auto-encoder-based VC system. We find that to leverage the U-Net architecture, a strong information bottleneck is necessary. The VQ-based method, which quantizes the latent vectors, can serve the purpose. The objective and the subjective evaluations show that the proposed method performs well in both audio naturalness and speaker similarity.
DOI: 10.1109/icassp40776.2020.9054734
发表时间: 2020-04
期刊: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Kaizhi Qian;Zeyu Jin;M. Hasegawa-Johnson;G. J. Mysore
通讯作者: Kaizhi Qian;Zeyu Jin;M. Hasegawa-Johnson;G. J. Mysore