Discrete Wavelet Denoising into MFCC for Noise Suppressive in Automatic Speech Recognition System

Discrete Wavelet Denoising into MFCC for Noise Suppressive in Automatic Speech Recognition System
复制标题

DOI:
10.22266/ijies2020.0430.08
复制
发表时间:
2020-04
影响因子:
--
通讯作者:
Hay Mar Soe Naing;Risanuri Hidayat;Rudy Hartanto;Y. Miyanaga
Hay Mar Soe Naing;Risanuri Hidayat;Rudy Hartanto;Y. Miyanaga
中科院分区:
--
文献类型:
--
作者:
Hay Mar Soe Naing;Risanuri Hidayat;Rudy Hartanto;Y. Miyanaga

文献摘要

相似文献

自动语音识别(ASR)是一项具有挑战性的任务,最大的问题是在存在背景噪声和语音的显著变异性的情况下。提取抗噪特征以适应由于噪声效应引起的语音质量下降一直是近年来研究的热点问题。本文提出了一种小波去噪方案的框架,并分析了不同的小波族和合适的阈值规则用于特征提取,以提高ASR系统的性能。语音识别器采用基于混合高斯模型的隐马尔可夫模型(GMM-HMM)和深度神经网络(DNN)-HMM。识别性能表明,在Aurora2数据库上,结合小波变换对Mel倒谱系数(MFCC)进行去噪,获得了较好的抗噪特征。采用Coiflet小波和里格苏尔阈值去噪的交叉熵DNN-HMM训练的正确率最高,10dB97.54%,5dB93.13%,0dB75.63%,−5dB37.29%。
: Automatic Speech Recognition (ASR) is a challenging task and the most problematic issues being in presence of background noise and substantial variability in speech. Extracting the noise-robust features adjust for speech degradations due to noise effect retained popular issue in recent years. This paper presented a framework for wavelet denoising scheme and analysed the different wavelet families and proper thresholding rule into feature extraction to enhance the performance of ASR system. Gaussian Mixture Model-based Hidden Markov Model (GMM-HMM) and Deep Neural Network (DNN)-HMM are used as the speech recognizer. The recognition performance shows that the noise-robust features are obtained while combining with the wavelet transform denoising into Mel Frequency Cepstral Coefficient (MFCC) on Aurora2 database. The best accuracy is gained by cross entropy DNN-HMM training using denoising with Coiflet wavelet and Rigrsure threshold, which provides 97.54% in 10dB, 93.13% in 5dB, 75.63% in 0dB and 37.29% in −5dB .