Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion
复制标题

DOI:
10.21437/interspeech.2020-2014
复制
发表时间:
2020-05
期刊:
The International Journal of Plant, Animal and Environmental Sciences
影响因子:
--
通讯作者:
Kun Zhou;Berrak Sisman;Mingyang Zhang;Haizhou Li
Kun Zhou;Berrak Sisman;Mingyang Zhang;Haizhou Li
中科院分区:
其他
文献类型:
--
作者:
Kun Zhou;Berrak Sisman;Mingyang Zhang;Haizhou Li

文献摘要

被引文献

相似文献

情感语音转换的目的是将语音的情感从一种状态转换到另一种状态,同时保持语言内容和说话人身份。以往的情感语音转换研究大多是在情感依赖于说话人的假设下进行的。我们认为,在口语中,说话者之间存在一个共同的情感表达代码,因此,说话者独立的情感状态之间的映射是可能的。在本文中,我们提出了一个独立于说话人的情感语音转换框架,可以转换任何人的情感,而不需要并行数据。我们提出了一种基于VAW-GAN的编码器-解码器结构来学习频谱和韵律映射。我们使用连续小波变换(CWT)进行韵律转换的时间依赖模型。我们还研究了使用F0作为解码器的额外输入,以提高情感转换性能。实验表明,所提出的说话人无关的框架实现了有竞争力的结果,为可见和不可见的发言人。
Emotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. The prior studies on emotional voice conversion are mostly carried out under the assumption that emotion is speaker-dependent. We consider that there is a common code between speakers for emotional expression in a spoken language, therefore, a speaker-independent mapping between emotional states is possible. In this paper, we propose a speaker-independent emotional voice conversion framework, that can convert anyone's emotion without the need for parallel data. We propose a VAW-GAN based encoder-decoder structure to learn the spectrum and prosody mapping. We perform prosody conversion by using continuous wavelet transform (CWT) to model the temporal dependencies. We also investigate the use of F0 as an additional input to the decoder to improve emotion conversion performance. Experiments show that the proposed speaker-independent framework achieves competitive results for both seen and unseen speakers.