Non-stationary feature extraction for automatic speech recognition

Non-stationary feature extraction for automatic speech recognition
复制标题

自动语音识别的非平稳特征提取

DOI:
10.1109/icassp.2011.5947530
复制
发表时间:
2011
期刊:
2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
F. Drepper
F. Drepper
中科院分区:
--
文献类型:
--
作者:
Zoltán Tüske;Pavel Golik;R. Schlüter;F. Drepper

文献摘要

被引文献

相似文献

在当前的语音识别系统中,主要使用基于短时傅里叶变换的特征,如MFCC。摒弃浊音短时平稳假设,将非平稳信号分析引入到ASR框架中。我们提出了一种基音自适应Gammatone滤波器组来提取新的声学特征。在Aurora 2和Aurora 4任务上证明了其噪声稳健性,其中所提出的特征优于标准MFCC。此外,通过Rover的成功组合实验表明了新功能与MFCC之间的差异。
In current speech recognition systems mainly Short-Time Fourier Transform based features like MFCC are applied. Dropping the short-time stationarity assumption of the voiced speech, this paper introduces the non-stationary signal analysis into the ASR framework. We present new acoustic features extracted by a pitch-adaptive Gammatone filter bank. The noise robustness was proved on AURORA 2 and 4 tasks, where the proposed features outperform the standard MFCC. Furthermore, successful combination experiments via ROVER indicate the differences between the new features and MFCC.