Distant-talking speech recognition based on a 3-D Viterbi search using a microphone array

Distant-talking speech recognition based on a 3-D Viterbi search using a microphone array
复制标题

使用麦克风阵列基于 3-D 维特比搜索的远距离通话语音识别

DOI:
10.1109/89.985542
复制
发表时间:
2002
期刊:
IEEE Transactions on Speech and Audio Processing
影响因子:
--
通讯作者:
K. Shikano
K. Shikano
中科院分区:
--
文献类型:
--
作者:
Takeshi Yamada;Satoshi Nakamura;K. Shikano

文献摘要

被引文献

相似文献

本文主要研究麦克风阵列在实际环境中实现远距离通话语音识别的方法。在远距离通话的情况下,用户可以在移动时在任意位置说话。因此,利用麦克风阵列进行高质量的语音采集,对说话人进行准确定位是非常重要的。然而,在嘈杂和混响的环境中定位移动的说话者是非常困难的。说话人定位错误导致语音识别性能下降。解决这一问题的一种方法是将语音识别过程和说话者本地化集成到一个统一的框架中。提出了一种新的基于三维维特比搜索的语音识别算法。三维维特比方法通过在每一帧中将波束引导到每个方向来提取参数向量的方向-时间序列,然后在由说话者方向、输入帧和HMM状态组成的三维网格空间中寻找最可能的路径。这意味着在统计框架内同时执行语音识别和说话者定位。为了评估三维维特比方法的性能,对真实环境数据进行了识别实验。实验结果表明,三维维特比方法显著提高了运动说话人和固定说话人情况下的识别性能。
This paper focuses on microphone arrays to realize distant-talking speech recognition in real environments. In distant-talking situations, users can speak at arbitrary positions while moving. Therefore, it,is very important for high quality speech acquisition using microphone arrays to localize a talker accurately. However, it is very difficult to localize a moving talker in noisy and reverberant environments. The talker localization errors result in performance degradation of speech recognition. One way to solve this problem is to integrate the speech recognition process and the talker localization into a unified framework. This paper proposes a new speech recognition algorithm based on a three-dimensional (3-D) Viterbi search. The 3-D Viterbi method extracts a direction-time sequence of parameter vectors by steering a beam to every direction in every frame, then finds the most likely path in a 3-D trellis space composed of talker directions, input frames and HMM states. This means that speech recognition and talker localization are performed simultaneously within a statistical framework. To evaluate the performance of the 3-D Viterbi method, recognition experiments for real environment data were carried out. The results confirmed that the 3-D Viterbi method drastically improves the recognition performance for the moving talker case as well as for the fixed-position talker case.