Incremental speaker adaptation with minimum error discriminative training for speaker identification

Incremental speaker adaptation with minimum error discriminative training for speaker identification
复制标题

用于说话人识别的最小误差判别训练的增量说话人适应

DOI:
10.1109/icslp.1996.607969
复制
发表时间:
1996
期刊:
Proceeding of Fourth International Conference on Spoken Language Processing. ICSLP '96
影响因子:
--
通讯作者:
L. Hernández
L. Hernández
中科院分区:
--
文献类型:
--
作者:
C. M. D. Alamo;J. Alvarez;C. D. L. Torre;F. J. Poyatos;L. Hernández

文献摘要

被引文献

相似文献

最小分类错误(MCE)已被证明是有效的,在提高说话人识别系统的性能。然而,仍然有问题需要解决,例如特定说话者的语音特征随时间的变化。在本文中,我们分析了退化的高斯混合模型(GMM)为基础的文本无关的说话人识别系统时,使用测试数据记录超过六个月后的训练会议,并试图避免这种退化,我们研究了使用监督自适应最大后验概率(MAP)估计和MCE的基础上。这些技术已被证明是语音识别中的说话人自适应提供了良好的效果。我们已经获得的主要结果是,通过从仅用来自会话1的语音训练的GMM模型开始,可以针对所有其他会话使用增量自适应来获得类似的识别结果,增量自适应使用每个说话者和会话仅2.5秒的语音作为MCE训练自适应过程的数据。我们还发现,在我们的极端实验设置中,MAP与MCE适应相结合时变得毫无帮助。
The minimum classification error (MCE) has been shown to be effective in improving the performance of a speaker identification system. However, there are still problems to solve, such as the variability of the voice characteristics of a particular speaker through time. In this paper, we analyze the degradation of a Gaussian mixture model (GMM) based text-independent speaker identification system when using test data recorded over six months after the training session, and, in an attempt to avoid this degradation, we study the use of supervised adaptation based on maximum a posteriori (MAP) estimation and MCE. These techniques have been shown to provide good results for speaker adaptation in speech recognition. The major result we have obtained is that, by starting with GMM models trained with only speech from session 1, similar identification results can be obtained for all the other sessions using an incremental adaptation using only 2.5 seconds of speech per speaker and session as data for the MCE training adaptation procedure. We have also found that, in our extreme experimental setup, MAP becomes unhelpful when combined with MCE adaptation.