Acoustic-to-articulatory inversion mapping with Gaussian mixture model

Acoustic-to-articulatory inversion mapping with Gaussian mixture model
复制标题

DOI:
10.21437/interspeech.2004-410
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
T. Toda;A. Black;K. Tokuda
T. Toda;A. Black;K. Tokuda
中科院分区:
其他
文献类型:
--
作者:
T. Toda;A. Black;K. Tokuda

文献摘要

被引文献

相似文献

本文介绍了使用高斯混合模型(GMM)的声学发音反演映射。声学参数和发音参数的对应关系由使用并行声学发音数据训练的GMM来建模。我们测量的性能基于GMM的映射,并调查使用多个声学帧作为输入功能,并使用多个混合的有效性。其结果是,它示出,虽然增加的混合物的数量是有用的,以减少估计误差,它会导致许多不连续的估计发音轨迹。为了解决这个问题,我们将考虑发音动态特征的最大似然估计(MLE)应用于基于GMM的映射。实验结果表明,使用动态特征的最大似然估计比基于高斯混合模型的映射应用低通滤波器平滑,可以估计更合适的发音运动。
This paper describes the acoustic-to-articulatory inversion mapping using a Gaussian Mixture Model (GMM). Correspondence of an acoustic parameter and an articulatory parameter is modeled by the GMM trained using the parallel acousticarticulatory data. We measure the performance of the GMMbased mapping and investigate the effectiveness of using multiple acoustic frames as an input feature and using multiple mixtures. As a result, it is shown that although increasing the number of mixtures is useful for reducing the estimation error, it causes many discontinuities in the estimated articulatory trajectories. In order to address this problem, we apply maximum likelihood estimation (MLE) considering articulatory dynamic features to the GMM-based mapping. Experimental results demonstrate that the MLE using dynamic features can estimate more appropriate articulatory movements compared with the GMM-based mapping applied smoothing by lowpass filter.