Mapping ultrasound-based articulatory images and vowel sounds with a deep neural network framework

Mapping ultrasound-based articulatory images and vowel sounds with a deep neural network framework
复制标题

使用深度神经网络框架映射基于超声的发音图像和元音声音

DOI:
10.1007/s11042-015-3038-y
复制
发表时间:
2016-05
影响因子:
3.6
通讯作者:
Dang, Jianwu
Dang, Jianwu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Zheng, Xinyuan;Lu, Wenhuan;He, Yuqing;Dang, Jianwu

文献摘要

参考文献

被引文献

相似文献

建立发音动作与对应语音的映射关系,对于语音障碍患者的语音训练和语音辅助工具的开发具有重要意义。在本文中,我们提出了一种新的深度学习框架,用于在发音信息和使用超声系统记录的同步语音之间创建双向映射。我们创建了一个包含六个中文元音的数据集,并采用基于受限玻尔兹曼机(RBM)的双峰深度自动编码器算法来学习舌头的语音和超声图像之间的相关性以及所获得的数据表示的权重矩阵。语音和超声图像,然后从提取的功能重建。我们的方法创建的超声图像的重建误差被发现是小于基于主成分分析(PCA)的方法。此外,重建的语音近似原始的平均共振峰误差(MFE)是小的。在使用基于RBM的深度自动编码器获取它们的共享表示之后,我们使用修改后的深度去噪自动编码器,使用深度神经网络(DNN)框架在舌头的超声图像和相应的声学信号之间进行映射。所获得的结果表明,我们所提出的方法的性能优于高斯混合模型(GMM)为基础的方法,它是比较。
Constructing a mapping between articulatory movements and corresponding speech could significantly facilitate speech training and the development of speech aids for voice disorder patients. In this paper, we propose a novel deep learning framework for the creation of a bidirectional mapping between articulatory information and synchronized speech recorded using an ultrasound system. We created a dataset comprising six Chinese vowels and employed the Bimodal Deep Autoencoders algorithm based on the Restricted Boltzmann Machine (RBM) to learn the correlation between speech and ultrasound images of the tongue and the weight matrices of the data representations obtained. Speech and ultrasound images were then reconstructed from the extracted features. The reconstruction error of the ultrasound images created with our method was found to be less than that of the approach based on Principal Components Analysis (PCA). Further, the reconstructed speech approximated the original as the mean formants error (MFE) was small. Following acquisition of their shared representations using the RBM-based deep autoencoder, we carried out mapping between ultrasound images of the tongue and corresponding acoustics signals with a Deep Neural Network (DNN) framework using the revised Deep Denoising Autoencoders. The results obtained indicate that the performance of our proposed method is better than that of a Gaussian Mixture Model (GMM)-based method to which it was compared.
DOI: 10.21437/interspeech.2006-213
发表时间: 2006-09
期刊: --
影响因子: --
作者:
Korin Richmond
通讯作者: Korin Richmond
DOI: 10.1109/icassp.2007.366989
发表时间: 2007-04
期刊: 2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP '07
影响因子: --
作者:
Karen Livescu;Ö. Çetin;M. Hasegawa-Johnson;Simon King;C. Bartels;Nash M. Borges;Arthur Kantor;Partha Lal;Lisa Yung;Ari Bezman;Stephen Dawson-Haggerty;B. Woods;Joe Frankel;M. Magimai-Doss;Kate Saenko
通讯作者: Karen Livescu;Ö. Çetin;M. Hasegawa-Johnson;Simon King;C. Bartels;Nash M. Borges;Arthur Kantor;Partha Lal;Lisa Yung;Ari Bezman;Stephen Dawson-Haggerty;B. Woods;Joe Frankel;M. Magimai-Doss;Kate Saenko
DOI: 10.1109/tsa.2003.822636
发表时间: 2004-04
期刊: IEEE Transactions on Speech and Audio Processing
影响因子: --
作者:
S. Hiroya;M. Honda
通讯作者: S. Hiroya;M. Honda
DOI: 10.1109/tmm.2006.888009
发表时间: 2007-04
影响因子: 7.3
作者:
Lei Xie;Zhi-Qiang Liu
通讯作者: Lei Xie;Zhi-Qiang Liu
DOI: 10.5555/1756006.1953039
发表时间: 2010-03
期刊: J. Mach. Learn. Res.
影响因子: --
作者:
Pascal Vincent;H. Larochelle;Isabelle Lajoie;Yoshua Bengio;Pierre-Antoine Manzagol
通讯作者: Pascal Vincent;H. Larochelle;Isabelle Lajoie;Yoshua Bengio;Pierre-Antoine Manzagol