Mel, linear, and antimel frequency cepstral coefficients in broad phonetic regions for telephone speaker recognition

Mel, linear, and antimel frequency cepstral coefficients in broad phonetic regions for telephone speaker recognition
复制标题

用于电话说话人识别的广泛语音区域中的梅尔、线性和反梅尔频率倒谱系数

DOI:
--
复制
发表时间:
2009
期刊:
Interspeech
影响因子:
--
通讯作者:
Eduardo López Gonzalo
Eduardo López Gonzalo
中科院分区:
--
文献类型:
--
作者:
Howard Lei;Eduardo López Gonzalo

文献摘要

被引文献

相似文献

We’ve examined the speaker discriminative power of mel-, antimeland linear-frequency cepstral coefficients (MFCCs, aMFCCs and LFCCs) in the nasal, vowel, and non-nasal consonant speech regions. Our inspiration came from the work of Lu and Dang in 2007, who showed that filterbank energies at some frequencies mainly outside the telephone bandwidth possess more speaker discriminative power due to physiological characteristics of speakers, and derived a set of cepstral coefficients that outperformed MFCCs in non-telephone speech. Using telephone speech, we’ve discovered that LFCCs gave 21.5% and 15.0% relative EER improvements over MFCCs in nasal and non-nasal consonant regions, agreeing with our filterbank energy f-ratio analysis. We’ve also found that using only the vowel region with MFCCs gives a 9.1% relative improvement over using all speech. Last, we’ve shown that a-MFCCs are valuable in combination, contributing to a system with 17.3% relative improvement over our baseline.