Recognizing the message and the messenger: biomimetic spectral analysis for robust speech and speaker recognition.

Recognizing the message and the messenger: biomimetic spectral analysis for robust speech and speaker recognition.
复制标题

DOI:
10.1007/s10772-012-9184-y
复制
发表时间:
2013
影响因子:
--
通讯作者:
Elhilali M
Elhilali M
中科院分区:
其他
文献类型:
--
作者:
Nemala SK;Patil K;Elhilali M

文献摘要

相似文献

人类非常善于在噪音中进行交流。然而,大多数语音处理系统(例如自动语音和说话人识别系统)在语音信号因看不见的背景失真而损坏时,性能会显着下降。拟议的工作探索了使用生物驱动的多分辨率频谱分析来表示语音。该方法侧重于语音信息丰富的频谱属性,并通过仔细选择模型参数来对语音信号进行复杂但计算高效的分析。此外,该方法利用对语音信号中的消息和说话人主导区域的信息论分析,并定义特征表示来解决语音和说话人识别等两个不同的任务。所提出的分析超越了标准梅尔倒谱系数 (MFCC) 及其增强变体(通过均值减法、方差归一化和时间序列过滤),并且在语音和说话人识别任务上比最先进的噪声鲁棒特征方案有了显着改进。
Humans are quite adept at communicating in presence of noise. However most speech processing systems, like automatic speech and speaker recognition systems, suffer from a significant drop in performance when speech signals are corrupted with unseen background distortions. The proposed work explores the use of a biologically-motivated multi-resolution spectral analysis for speech representation. This approach focuses on the information-rich spectral attributes of speech and presents an intricate yet computationally-efficient analysis of the speech signal by careful choice of model parameters. Further, the approach takes advantage of an information-theoretic analysis of the message and speaker dominant regions in the speech signal, and defines feature representations to address two diverse tasks such as speech and speaker recognition. The proposed analysis surpasses the standard Mel-Frequency Cepstral Coefficients (MFCC), and its enhanced variants (via mean subtraction, variance normalization and time sequence filtering) and yields significant improvements over a state-of-the-art noise robust feature scheme, on both speech and speaker recognition tasks.