Classification of fricative consonants for speech enhancement in hearing devices.

Classification of fricative consonants for speech enhancement in hearing devices.
复制标题

DOI:
10.1371/journal.pone.0095001
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Kokkinakis K
Kokkinakis K
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kong YY;Mullangi A;Kokkinakis K

文献摘要

参考文献

被引文献

相似文献

探讨三组不同发音位置的摩擦辅音的声学特征和分类方法。采用支持向量机(SVM)算法对从TIMIT数据库中提取的摩擦音在不同信噪比下的安静噪声和语音杂音噪声中进行分类。光谱特征包括四谱矩、峰值、斜率、Mel-frequency倒谱系数(MFCC)、Gammatone滤波器输出和快速傅里叶变换(FFT)频谱的幅度。分析框架被限制为只有8毫秒。此外,研究了将高维特征向量投影到低维空间的常用线性和非线性主成分分析降维技术。13个MFCC系数,14或24个Gammatone滤波器输出,在安静和+10 dB信噪比下的分类性能大于或等于85%。使用14个高于1 kHz的Gammatone滤波器输出,在+20到+5 dB信噪比的宽信噪比范围内,分类精度仍然很高(大于80%)。仅使用从短时间窗口提取的频谱特征就可以实现高水平的静音和噪声摩擦辅音分类精度。这项工作的结果对听力设备语音增强算法的发展有直接的影响。
To investigate a set of acoustic features and classification methods for the classification of three groups of fricative consonants differing in place of articulation. A support vector machine (SVM) algorithm was used to classify the fricatives extracted from the TIMIT database in quiet and also in speech babble noise at various signal-to-noise ratios (SNRs). Spectral features including four spectral moments, peak, slope, Mel-frequency cepstral coefficients (MFCC), Gammatone filters outputs, and magnitudes of fast Fourier Transform (FFT) spectrum were used for the classification. The analysis frame was restricted to only 8 msec. In addition, commonly-used linear and nonlinear principal component analysis dimensionality reduction techniques that project a high-dimensional feature vector onto a lower dimensional space were examined. With 13 MFCC coefficients, 14 or 24 Gammatone filter outputs, classification performance was greater than or equal to 85% in quiet and at +10 dB SNR. Using 14 Gammatone filter outputs above 1 kHz, classification accuracy remained high (greater than 80%) for a wide range of SNRs from +20 to +5 dB SNR. High levels of classification accuracy for fricative consonants in quiet and in noise could be achieved using only spectral features extracted from a short time window. Results of this work have a direct impact on the development of speech enhancement algorithms for hearing devices.
DOI: 10.1016/j.specom.2011.07.008
发表时间: 2012-01-01
影响因子: 3.2
作者:
Kong, Ying-Yee;Mullangi, Ala
通讯作者: Mullangi, Ala
DOI: 10.1121/1.1357814
发表时间: 2001-05-01
影响因子: 2.4
作者:
Ali, AMA;Van der Spiegel, J;Mueller, P
通讯作者: Mueller, P
DOI: 10.1121/1.410152
发表时间: 1994-10-01
影响因子: 2.4
作者:
BYRNE, D;DILLON, H;LUDVIGSEN, C
通讯作者: LUDVIGSEN, C
DOI: 10.1080/14992020601188591
发表时间: 2007-06-01
影响因子: 2.7
作者:
Robinson, Joanna D.;Baer, Thomas;Moore, Brian C. J.
通讯作者: Moore, Brian C. J.
DOI: 10.1162/089976698300017467
发表时间: 1998-07-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Scholkopf, B;Smola, A;Muller, KR
通讯作者: Muller, KR