One-model speech recognition and synthesis based on articulatory movement HMMs

One-model speech recognition and synthesis based on articulatory movement HMMs
复制标题

DOI:
10.21437/interspeech.2010-31
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
T. Nitta;Takayuki Onoda;Masashi Kimura;Y. Iribe;Kouichi Katsurada
T. Nitta;Takayuki Onoda;Masashi Kimura;Y. Iribe;Kouichi Katsurada
中科院分区:
其他
文献类型:
--
作者:
T. Nitta;Takayuki Onoda;Masashi Kimura;Y. Iribe;Kouichi Katsurada

文献摘要

相似文献

本文描述了基于公共发音运动模型的单模型语音识别(SR)和语音合成(SS)。SR引擎具有发音特征(AF)提取器和对发音姿势进行建模的基于HMM的分类器。音素识别的实验结果表明,即使训练数据仅限于单个说话人,AF也优于MFCC。在SS引擎中,应用相同的说话人不变HMM来生成AF序列,然后,在将AF转换为声道参数之后,通过PARCOR滤波器合成语音信号以及残余信号。还描述了使用AF交换的音素到音素语音转换。
One-model speech recognition (SR) and speech synthesis (SS) based on a common articulatory movement model are described herein. The SR engine has an articulatory feature (AF) extractor and an HMM based classifier that models articulatory gestures. Experimental results of a phoneme recognition task show that the AF outperforms MFCC even if the training data are limited to a single speaker. In the SS engine, the same speaker-invariant HMM is applied to generate an AF sequence, and then, after converting AFs into vocal tract parameters, a speech signal is synthesized by a PARCOR filter, together with a residual signal. Phoneme-to-phoneme speech conversion, using AF exchange, is also described.