Mapping ultrasound-based articulatory images and vowel sounds with a deep neural network framework
Mapping ultrasound-based articulatory images and vowel sounds with a deep neural network framework
复制标题
使用深度神经网络框架映射基于超声的发音图像和元音声音
DOI:
10.1007/s11042-015-3038-y
复制
发表时间:
2016-05
影响因子:
3.6
通讯作者:
Dang, Jianwu
中科院分区:
文献类型:
--
作者:
Zheng, Xinyuan;Lu, Wenhuan;He, Yuqing;Dang, Jianwu
Constructing a mapping between articulatory movements and corresponding speech could significantly facilitate speech training and the development of speech aids for voice disorder patients. In this paper, we propose a novel deep learning framework for the creation of a bidirectional mapping between articulatory information and synchronized speech recorded using an ultrasound system. We created a dataset comprising six Chinese vowels and employed the Bimodal Deep Autoencoders algorithm based on the Restricted Boltzmann Machine (RBM) to learn the correlation between speech and ultrasound images of the tongue and the weight matrices of the data representations obtained. Speech and ultrasound images were then reconstructed from the extracted features. The reconstruction error of the ultrasound images created with our method was found to be less than that of the approach based on Principal Components Analysis (PCA). Further, the reconstructed speech approximated the original as the mean formants error (MFE) was small. Following acquisition of their shared representations using the RBM-based deep autoencoder, we carried out mapping between ultrasound images of the tongue and corresponding acoustics signals with a Deep Neural Network (DNN) framework using the revised Deep Denoising Autoencoders. The results obtained indicate that the performance of our proposed method is better than that of a Gaussian Mixture Model (GMM)-based method to which it was compared.
登录
查看更多内容
DOI:
10.21437/interspeech.2006-213
发表时间:
2006-09
期刊:
--
影响因子:
--
作者:
Korin Richmond
通讯作者:
Korin Richmond
DOI:
10.1109/icassp.2007.366989
发表时间:
2007-04
期刊:
2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP '07
影响因子:
--
作者:
Karen Livescu;Ö. Çetin;M. Hasegawa-Johnson;Simon King;C. Bartels;Nash M. Borges;Arthur Kantor;Partha Lal;Lisa Yung;Ari Bezman;Stephen Dawson-Haggerty;B. Woods;Joe Frankel;M. Magimai-Doss;Kate Saenko
通讯作者:
Karen Livescu;Ö. Çetin;M. Hasegawa-Johnson;Simon King;C. Bartels;Nash M. Borges;Arthur Kantor;Partha Lal;Lisa Yung;Ari Bezman;Stephen Dawson-Haggerty;B. Woods;Joe Frankel;M. Magimai-Doss;Kate Saenko
DOI:
10.1109/tsa.2003.822636
发表时间:
2004-04
期刊:
IEEE Transactions on Speech and Audio Processing
影响因子:
--
作者:
S. Hiroya;M. Honda
通讯作者:
S. Hiroya;M. Honda
影响因子:
7.3
作者:
Lei Xie;Zhi-Qiang Liu
通讯作者:
Lei Xie;Zhi-Qiang Liu
DOI:
10.5555/1756006.1953039
发表时间:
2010-03
期刊:
J. Mach. Learn. Res.
影响因子:
--
作者:
Pascal Vincent;H. Larochelle;Isabelle Lajoie;Yoshua Bengio;Pierre-Antoine Manzagol
通讯作者:
Pascal Vincent;H. Larochelle;Isabelle Lajoie;Yoshua Bengio;Pierre-Antoine Manzagol