Probabilistic integration of joint density model and speaker model for voice conversion

Probabilistic integration of joint density model and speaker model for voice conversion
复制标题

DOI:
10.21437/interspeech.2010-496
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
D. Saito;Shinji Watanabe;Atsushi Nakamura;N. Minematsu
D. Saito;Shinji Watanabe;Atsushi Nakamura;N. Minematsu
中科院分区:
其他
文献类型:
--
作者:
D. Saito;Shinji Watanabe;Atsushi Nakamura;N. Minematsu

文献摘要

被引文献

相似文献

本文描述了一种使用联合密度模型和说话人模型的语音转换新方法。在语音转换研究中,基于具有源和目标说话者的联合向量的概率密度的高斯混合模型(GMM)的方法被广泛用于估计转换。然而,为了获得足够的质量,它们需要一个平行语料库,其中包含两个说话者所说的大量具有相同语言内容的话语。此外,当训练数据量较小时,联合密度 GMM 方法常常会受到过度训练的影响。为了弥补这些问题,我们提出了一种新方法,使用概率公式将目标的说话者 GMM 与联合密度模型集成。该方法独立地训练具有少量并行话语的联合密度模型和具有目标的非并行数据的说话人模型。它减轻了源扬声器的负担。实验证明了该方法的有效性,特别是当平行语料量较小时。索引术语:语音转换、联合密度模型、说话人模型、概率统一
This paper describes a novel approach to voice conversion using both a joint density model and a speaker model. In voice conversion studies, approaches based on Gaussian Mixture Model (GMM) with probabilistic densities of joint vectors of a source and a target speakers are widely used to estimate a transformation. However, for sufficient quality, they require a parallel corpus which contains plenty of utterances with the same linguistic content spoken by both the speakers. In addition, the joint density GMM methods often suffer from over-training effects when the amount of training data is small. To compensate for these problems, we propose a novel approach to integrate the speaker GMM of the target with the joint density model using probabilistic formulation. The proposed method trains the joint density model with a few parallel utterances, and the speaker model with non-parallel data of the target, independently. It eases the burden on the source speaker. Experiments demonstrate the effectiveness of the proposed method, especially when the amount of the parallel corpus is small. Index Terms: voice conversion, joint density model, speaker model, probabilistic unification