A Multistream Feature Framework Based on Bandpass Modulation Filtering for Robust Speech Recognition.

A Multistream Feature Framework Based on Bandpass Modulation Filtering for Robust Speech Recognition.
复制标题

DOI:
10.1109/tasl.2012.2219526
复制
发表时间:
2013-03
期刊:
IEEE transactions on audio, speech, and language processing
影响因子:
--
通讯作者:
Elhilali M
Elhilali M
中科院分区:
其他
文献类型:
--
作者:
Nemala SK;Patil K;Elhilali M

文献摘要

被引文献

相似文献

有强有力的神经生理学证据表明,大脑中的语音信号处理是沿着沿着的路径发生的,这些路径将互补信息编码在信号中。这些并行流是围绕着慢与快的二元性组织的:在频谱和时间维度上,粗信号动态似乎与快速变化的调制分开处理。我们适应这样的二元性,在一个强大的说话人独立的音素识别的多流框架。这里提出的方案围绕着一个多径带通调制分析的语音声音与每个流覆盖整个范围的时间和频谱调制。通过执行带通操作沿着的频谱和时间维度,该方案避免了经典的特征爆炸问题,以前的多流方法,同时保持并行性和本地化的特征分析的优势。所提出的架构的结果在显着改善标准和国家的最先进的音素识别的功能计划,特别是在存在非平稳噪声,混响和信道失真。
There is strong neurophysiological evidence suggesting that processing of speech signals in the brain happens along parallel paths which encode complementary information in the signal. These parallel streams are organized around a duality of slow vs. fast: Coarse signal dynamics appear to be processed separately from rapidly changing modulations both in the spectral and temporal dimensions. We adapt such duality in a multistream framework for robust speaker-independent phoneme recognition. The scheme presented here centers around a multi-path bandpass modulation analysis of speech sounds with each stream covering an entire range of temporal and spectral modulations. By performing bandpass operations along the spectral and temporal dimensions, the proposed scheme avoids the classic feature explosion problem of previous multistream approaches while maintaining the advantage of parallelism and localized feature analysis. The proposed architecture results in substantial improvements over standard and state-of-the-art feature schemes for phoneme recognition, particularly in presence of nonstationary noise, reverberation and channel distortions.