Vaw-Gan For Disentanglement And Recomposition Of Emotional Elements In Speech

Vaw-Gan For Disentanglement And Recomposition Of Emotional Elements In Speech
复制标题

DOI:
10.1109/slt48900.2021.9383526
复制
发表时间:
2020-11
期刊:
2021 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
Kun Zhou;Berrak Sisman;Haizhou Li
Kun Zhou;Berrak Sisman;Haizhou Li
中科院分区:
其他
文献类型:
--
作者:
Kun Zhou;Berrak Sisman;Haizhou Li

文献摘要

被引文献

相似文献

情感语音转换(EVC)的目的是将语音的情感从一种状态转换到另一种状态,同时保持语言内容和说话人身份。本文利用变分自编码Wasserstein生成对抗网络(VAW-GAN)研究语音中情感成分的分解和重组。我们提出了一种基于VAW-GAN的说话人相关EVC框架,该框架包括两个VAW-GAN管道,一个用于频谱转换,另一个用于韵律转换。我们训练了一个频谱编码器,从频谱特征中分离情感和韵律(F0)信息;我们还训练了一个韵律编码器,从语言韵律中分离韵律的情感调制(情感韵律)。在运行时,频谱VAW-GAN的解码器以韵律VAW-GAN的输出为条件。声码器采用转换后的频谱和韵律特征来生成目标情感语音。实验验证了我们提出的方法在客观和主观评价的有效性。
Emotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition of emotional elements in speech through variational autoencoding Wasserstein generative adversarial network (VAW-GAN). We propose a speaker-dependent EVC framework based on VAW-GAN, that includes two VAW-GAN pipelines, one for spectrum conversion, and another for prosody conversion. We train a spectral encoder that disentangles emotion and prosody (F0) information from spectral features; we also train a prosodic encoder that disentangles emotion modulation of prosody (affective prosody) from linguistic prosody. At run-time, the decoder of spectral VAW-GAN is conditioned on the output of prosodic VAW-GAN. The vocoder takes the converted spectral and prosodic features to generate the target emotional speech. Experiments validate the effectiveness of our proposed method in both objective and subjective evaluations.