COMPARISON OF PARAMETRIC REPRESENTATIONS FOR MONOSYLLABIC WORD RECOGNITION IN CONTINUOUSLY SPOKEN SENTENCES

COMPARISON OF PARAMETRIC REPRESENTATIONS FOR MONOSYLLABIC WORD RECOGNITION IN CONTINUOUSLY SPOKEN SENTENCES
复制标题

DOI:
10.1109/tassp.1980.1163420
复制
发表时间:
1980-01-01
期刊:
IEEE TRANSACTIONS ON ACOUSTICS SPEECH AND SIGNAL PROCESSING
影响因子:
--
通讯作者:
MERMELSTEIN, P
MERMELSTEIN, P
中科院分区:
其他
文献类型:
--
作者:
DAVIS, SB;MERMELSTEIN, P

文献摘要

被引文献

相似文献

在面向音节的连续语音识别系统中,比较了声信号的几种参数表示对单词识别性能的影响。词汇包括许多语音相似的单音节单词,因此重点是在面对句法和持续时间变化时保留语音上重要的声学信息的能力。对于每个参数集(基于mel-frequency倒谱、线性频率倒谱、线性预测倒谱、线性预测谱或一组反射系数),使用有效的动态扭曲方法生成单词模板,并将测试数据与模板进行时间注册。每6.4 ms计算一组10个mel频率倒谱系数产生了最佳性能,即对两个说话者中的每个人的识别率分别为96.5%和95.0%。梅尔频率倒谱系数的优越性能可能归因于这样一个事实,即它们更好地代表了短期语音频谱的感知相关方面。
Several parametric representations of the acoustic signal were compared with regard to word recognition performance in a syllable-oriented continuous speech recognition system. The vocabulary included many phonetically similar monosyllabic words, therefore the emphasis was on the ability to retain phonetically significant acoustic information in the face of syntactic and duration variations. For each parameter set (based on a mel-frequency cepstrum, a linear frequency cepstrum, a linear prediction cepstrum, a linear prediction spectrum, or a set of reflection coefficients), word templates were generated using an efficient dynamic warping method, and test data were time registered with the templates. A set of ten mel-frequency cepstrum coefficients computed every 6.4 ms resulted in the best performance, namely 96.5 percent and 95.0 percent recognition with each of two speakers. The superior performance of the mel-frequency cepstrum coefficients may be attributed to the fact that they better represent the perceptually relevant aspects of the short-term speech spectrum.