A new method for speaker adaptation using bilinear model

A new method for speaker adaptation using bilinear model
复制标题

DOI:
10.1109/icassp.2009.4960596
复制
发表时间:
2009-04
期刊:
2009 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
H. Song;Yongwon Jeong;H. S. Kim
H. Song;Yongwon Jeong;H. S. Kim
中科院分区:
其他
文献类型:
--
作者:
H. Song;Yongwon Jeong;H. S. Kim

文献摘要

被引文献

相似文献

提出了一种基于双线性模型的说话人自适应方法。双线性模型可以在训练数据库中独立地表达说话人的特征(风格)和跨说话人的音素(内容)。利用与说话人和音素空间无关的双线性映射矩阵实现了从每个说话人和音素空间到观测空间的映射。我们将双线性模型应用于说话人自适应。利用新说话人的自适应数据,通过估计风格(说话人)特定矩阵来建立说话人自适应模型。实验结果表明,该方法优于本征语音和MLLR。在用于说话人自适应的词汇无关孤立词识别中,与使用50个词进行自适应的本征语音和MLLR相比,双线性模型分别降低了约38%和约10%的单词错误率。
In this paper, a novel method for speaker adaptation using bilinear model is proposed. Bilinear model can express both characteristics of speakers (style) and phonemes across speakers (content) independently in a training database. The mapping from each speaker and phoneme space to observation space is carried out using bilinear mapping matrix which is independent of speaker and phoneme space. We apply the bilinear model to speaker adaption. Using adaptation data from a new speaker, speaker-adapted model is built by estimating the style(speaker)-specific matrix. Experimental results showed that the proposed method outperformed eigenvoice and MLLR. In vocabulary-independent isolated word recognition for speaker adaptation, bilinear model reduced word error rate by about 38% and about 10% compared to eigenvoice and MLLR respectively using 50 words for adaptation.