Noise robust speech recognition using F0 contour extracted by hough transform

Noise robust speech recognition using F0 contour extracted by hough transform
复制标题

使用霍夫变换提取的 F0 轮廓进行噪声鲁棒语音识别

DOI:
10.21437/icslp.2002-313
复制
发表时间:
2002
期刊:
--
影响因子:
--
通讯作者:
S. Furui
S. Furui
中科院分区:
--
文献类型:
--
作者:
K. Iwano;Takahiro Seki;S. Furui

文献摘要

被引文献

相似文献

提出了一种利用韵律信息的抗噪语音识别方法。在日语中,基频(F0)轮廓代表短语语调和单词重音信息。因此,它传达了韵律短语和单词边界的信息。本文首先提出了一种基于Hough变换的噪声鲁棒F0提取方法,该方法在各种噪声环境下都能获得较高的提取率。然后提出了一种基于音节H_2的语音识别方法,该方法同时利用了语音的分段谱特征和F_0轮廓。在不同的噪声和信噪比条件下,对11名男性说话人发出的连续数字进行了非特定人实验。在所有噪声条件下,识别准确率都得到了提高,数字准确率的最佳绝对提高约为4.7%。这种改进是由于更精确的数字边界检测的鲁棒韵律信息。
This paper proposes a noise robust speech recognition method using prosodic information. In Japanese, fundamental frequency (F0) contour represents phrase intonation and word accent information. Consequently, it conveys information about prosodic phrase and word boundaries. This paper first proposes a noise robust F0 extraction method using Hough transform, which achieves high extraction rates under various noise environments. Then it proposes a robust speech recognition method using syllable HMMs which model both segmental spectral features and F0 contours. Speaker-independent experiments are conducted using connected digits uttered by 11 male speakers in various kinds of noise and SNR conditions. The recognition accuracy is improved in all noise conditions, and the best absolute improvement of digit accuracy is about 4.7%. This improvement is achieved due to the more precise digit boundary detection by the robust prosodic information.