A Diffeomorphic Flow-Based Variational Framework for Multi-Speaker Emotion Conversion

A Diffeomorphic Flow-Based Variational Framework for Multi-Speaker Emotion Conversion
复制标题

DOI:
10.1109/taslp.2022.3209948
复制
发表时间:
2022-11
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Ravi Shankar;Hsi-Wei Hsieh;N. Charon;A. Venkataraman
Ravi Shankar;Hsi-Wei Hsieh;N. Charon;A. Venkataraman
中科院分区:
其他
文献类型:
--
作者:
Ravi Shankar;Hsi-Wei Hsieh;N. Charon;A. Venkataraman

文献摘要

相似文献

本文介绍了一种新的框架,在语音中的非平行情感转换。我们的框架基于两个关键贡献。首先,我们提出了流行的Cycle-GAN模型的随机版本。我们修改后的损失函数引入了Kullback-Leibler(KL)发散项,该项将生成器学习的源和目标数据分布对齐,从而克服了样本生成的局限性。通过使用变分近似这个随机损失函数,我们表明,我们的KL发散项可以通过一对密度泛函来实现。我们将这种新的架构称为变分循环GAN(VCGAN)。其次,我们将目标情感的韵律特征建模为源韵律特征的平滑和可学习的变形。这种方法提供了隐式正则化,其在与看不见的和分布外的扬声器的更好的范围对准方面提供了关键优势。我们进行了严格的实验和比较研究,以证明我们提出的框架是相当强大的高性能对几个国家的最先进的基线。
This paper introduces a new framework for non-parallel emotion conversion in speech. Our framework is based on two key contributions. First, we propose a stochastic version of the popular Cycle-GAN model. Our modified loss function introduces a Kullback–Leibler (KL) divergence term that aligns the source and target data distributions learned by the generators, thus overcoming the limitations of sample-wise generation. By using a variational approximation to this stochastic loss function, we show that our KL divergence term can be implemented via a paired density discriminator. We term this new architecture a variational Cycle-GAN (VCGAN). Second, we model the prosodic features of target emotion as a smooth and learnable deformation of the source prosodic features. This approach provides implicit regularization that offers key advantages in terms of better range alignment to unseen and out-of-distribution speakers. We conduct rigorous experiments and comparative studies to demonstrate that our proposed framework is fairly robust with high performance against several state-of-the-art baselines.