Distant-talking accent recognition by combining GMM and DNN

Distant-talking accent recognition by combining GMM and DNN
复制标题

DOI:
10.1007/s11042-015-2935-4
复制
发表时间:
2015-09
影响因子:
3.6
通讯作者:
Khomdet Phapatanaburi;Longbiao Wang;Ryota Sakagami;Zhaofeng Zhang;Ximin Li;M. Iwahashi
Khomdet Phapatanaburi;Longbiao Wang;Ryota Sakagami;Zhaofeng Zhang;Ximin Li;M. Iwahashi
中科院分区:
计算机科学4区
文献类型:
--
作者:
Khomdet Phapatanaburi;Longbiao Wang;Ryota Sakagami;Zhaofeng Zhang;Ximin Li;M. Iwahashi

文献摘要

被引文献

相似文献

近年来,自动口音识别越来越受到人们的关注。然而,很少有研究关注远距离通话环境中的口音识别,这对于提高非母语口音的远距离通话语音识别性能非常重要。在本文中,我们应用高斯混合模型(GMM)和深度神经网络(DNN)来识别混响环境中的说话者口音。还提出了可能性与这两种方法的结合。在混响环境中,口音识别率从 GMM 的 90.7% 提高到 DNN 的 93.0%。 GMM 和 DNN 的组合实现了 97.5% 的识别率,由于 GMM 和 DNN 的互补,其识别率优于单独的 GMM 和 DNN。相对误差分别比基于 GMM 的方法减少了 73.1%,比基于 DNN 的方法减少了 64.3%。
Recently, automatic accent recognition has been paid more and more attentions. However, there are few researches focusing on accent recognition in distant-talking environment which is very important for improving distant-talking speech recognition performance with non-native accents. In this paper, we apply Gaussian Mixture Models (GMM) and Deep Neural Network (DNN) to identify the speaker accent in reverberant environments. The combination of likelihood with these two approaches is also proposed. In reverberant environment, the accent recognition rate was improved from 90.7 % with GMM to 93.0 % with DNN. The combination of GMM and DNN achieved recognition rate of 97.5 %, which outperformed than the individual GMM and DNN because the complementation of GMM and DNN. The relative error reduction is 73.1 % than the GMM-based method and 64.3 % than the DNN-based method, respectively.