Modulation spectrum-based post-filter for GMM-based Voice Conversion

Modulation spectrum-based post-filter for GMM-based Voice Conversion
复制标题

DOI:
10.1109/apsipa.2014.7041540
复制
发表时间:
2014-12
期刊:
Signal and Information Processing Association Annual Summit and Conference (APSIPA), 2014 Asia-Pacific
影响因子:
--
通讯作者:
Shinnosuke Takamichi;T. Toda;A. Black;Satoshi Nakamura
Shinnosuke Takamichi;T. Toda;A. Black;Satoshi Nakamura
中科院分区:
其他
文献类型:
--
作者:
Shinnosuke Takamichi;T. Toda;A. Black;Satoshi Nakamura

文献摘要

被引文献

相似文献

针对基于高斯混合模型(GMM)的语音转换(VC)中存在的过平滑效应进行了研究。统计方法的灵活使用是这种方法被广泛应用于基于语音的系统的主要原因之一。然而,过平滑的语音参数转换质量下降是统计建模不可避免的问题。在转换步骤中解决这种过度平滑的常见方法之一是补偿生成的特征,例如全局方差(GV),其明确表示过度平滑效果。在统计的文语转换(TTS)合成中,我们引入了调制谱(MS),它是GV的一种扩展形式,并在基于隐马尔可夫模型(HMM)的TTS合成中提出了基于MS的后置滤波器(MSPF)。在本文中,我们将MSPF应用到基于GMM的VC。由于语音参数的MS通过基于GMM的转换过程被降级,由于MS修改转换后的参数,我们执行后滤波器。实验评估产生的质量效益所提出的后置滤波器。
This paper addresses an over-smoothing effect in Gaussian Mixture Model (GMM)-based Voice Conversion (VC). The flexible use of the statistical approach is one of the major reason why this approach is widely applied to the speech-based systems. However, quality degradation by over-smoothed speech parameter converted is unavoidable problem of statistical modeling. One of common approaches to this over-smoothness in conversion step is to compensate generated features, such as Global Variance (GV), that explicitly express the over-smoothing effect. In statistical Text-To-Speech (TTS) synthesis, we have recently introduced a Modulation Spectrum (MS) which is an extended form of GV, and have proposed MS-based Post-Filter (MSPF) in Hidden Markov Model (HMM)-based TTS synthesis. In this paper, we apply the MSPF to GMM-based VC. Because the MS of speech parameters is degraded through GMM-based conversion process, we perform the post-filter due to MS modification of converted parameters. The experimental evaluation yields the quality benefits by the proposed post-filter.