Crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder

Crank: An Open-Source Software for Nonparallel Voice Conversion Based on Vector-Quantized Variational Autoencoder
复制标题

DOI:
10.1109/icassp39728.2021.9413959
复制
发表时间:
2021-03
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Kazuhiro Kobayashi;Wen-Chin Huang;Yi-Chiao Wu;Patrick Lumban Tobing;Tomoki Hayashi;T. Toda
Kazuhiro Kobayashi;Wen-Chin Huang;Yi-Chiao Wu;Patrick Lumban Tobing;Tomoki Hayashi;T. Toda
中科院分区:
其他
文献类型:
--
作者:
Kazuhiro Kobayashi;Wen-Chin Huang;Yi-Chiao Wu;Patrick Lumban Tobing;Tomoki Hayashi;T. Toda

文献摘要

被引文献

相似文献

在本文中,我们提出了一种用于开发非并行语音转换(VC)系统的开源软件,名为crank。尽管我们在上一届 VC 挑战赛中发布了基于高斯混合模型的开源 VC 软件 sprocket,但应用任何语音语料库都不是一件容易的事,因为需要准备源说话人和目标说话人的并行话语来建模统计转换函数。为了解决这个问题,在本研究中,我们开发了一种新的开源 VC 软件,使用户能够仅使用非并行语音语料库来对转换函数进行建模。为了实现 VC 软件,我们使用了矢量量化变分自动编码器 (VQVAE)。为了快速检验该研究领域开发的最新技术的有效性,Crank还支持基于自动编码器的VC方法的几项代表性工作,例如分层架构、循环架构、生成对抗网络、说话人对抗训练和神经声码器的使用。此外,还可以基于 MOSNet 自动估计客观度量,例如梅尔倒谱失真和伪平均意见得分。在本文中,我们描述了曲柄中开发的代表性功能,并通过客观评价进行了简要比较。
In this paper, we present an open-source software for developing a nonparallel voice conversion (VC) system named crank. Although we have released an open-source VC software based on the Gaussian mixture model named sprocket in the last VC Challenge, it is not straightforward to apply any speech corpus because it is necessary to prepare parallel utterances of source and target speakers to model a statistical conversion function. To address this issue, in this study, we developed a new open-source VC software that enables users to model the conversion function by using only a nonparallel speech corpus. For implementing the VC software, we used a vector-quantized variational autoencoder (VQVAE). To rapidly examine the effectiveness of recent technologies developed in this research field, crank also supports several representative works for autoencoder-based VC methods such as the use of hierarchical architectures, cyclic architectures, generative adversarial networks, speaker adversarial training, and neural vocoders. Moreover, it is possible to automatically estimate objective measures such as mel-cepstrum distortion and pseudo mean opinion score based on MOSNet. In this paper, we describe representative functions developed in crank and make brief comparisons by objective evaluations.