One-model speech recognition and synthesis based on articulatory movement HMMs
One-model speech recognition and synthesis based on articulatory movement HMMs
复制标题
DOI:
10.21437/interspeech.2010-31
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
T. Nitta;Takayuki Onoda;Masashi Kimura;Y. Iribe;Kouichi Katsurada
中科院分区:
文献类型:
--
作者:
T. Nitta;Takayuki Onoda;Masashi Kimura;Y. Iribe;Kouichi Katsurada
One-model speech recognition (SR) and speech synthesis (SS) based on a common articulatory movement model are described herein. The SR engine has an articulatory feature (AF) extractor and an HMM based classifier that models articulatory gestures. Experimental results of a phoneme recognition task show that the AF outperforms MFCC even if the training data are limited to a single speaker. In the SS engine, the same speaker-invariant HMM is applied to generate an AF sequence, and then, after converting AFs into vocal tract parameters, a speech signal is synthesized by a PARCOR filter, together with a residual signal. Phoneme-to-phoneme speech conversion, using AF exchange, is also described.