Microscopic prediction of speech recognition for listeners with normal hearing in noise using an auditory model

Microscopic prediction of speech recognition for listeners with normal hearing in noise using an auditory model
复制标题

DOI:
10.1121/1.3224721
复制
发表时间:
2009-11-01
影响因子:
2.4
通讯作者:
Brand, Thomas
Brand, Thomas
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Juergens, Tim;Brand, Thomas

文献摘要

被引文献

相似文献

本研究比较了语音识别的微观模型的语音形状噪声的音素识别性能与正常听力的听众的性能。“微观”在这个模型中被定义为双重的。首先,语音识别率预测的音素的音素的基础上。其次,微观建模是指通过模仿人类听觉处理的基本部分来处理待识别的信号波形。该模型基于Holube和Kollmeier的方法[J. Acoust.美国社会100,1703-1716(1996)],并且由心理声学和生理动机的预处理和简单的动态时间扭曲语音识别器组成。该模型进行评估,同时提出无意义的讲话在一个封闭的范例。平均音素识别率,具体音素识别率,和音素混淆进行了分析。不同的感知距离的措施和模型的先验知识的影响进行了研究。结果表明,人的表现可以预测该模型使用最佳检测器,即,用于识别器训练和测试的相同语音波形。最好的模型性能产生的距离措施,主要集中在小的感知距离和忽略离群值。(C)2009年,美国声学学会。[DOI:10.1121/1.3224721]
This study compares the phoneme recognition performance in speech-shaped noise of a microscopic model for speech recognition with the performance of normal-hearing listeners. "Microscopic" is defined in terms of this model twofold. First, the speech recognition rate is predicted on a phoneme-by-phoneme basis. Second, microscopic modeling means that the signal waveforms to be recognized are processed by mimicking elementary parts of human's auditory processing. The model is based on an approach by Holube and Kollmeier [J. Acoust. Soc. Am. 100, 1703-1716 (1996)] and consists of a psychoacoustically and physiologically motivated preprocessing and a simple dynamic-time-warp speech recognizer. The model is evaluated while presenting nonsense speech in a closed-set paradigm. Averaged phoneme recognition rates, specific phoneme recognition rates, and phoneme confusions are analyzed. The influence of different perceptual distance measures and of the model's a-priori knowledge is investigated. The results show that human performance can be predicted by this model using an optimal detector, i.e., identical speech waveforms for both training of the recognizer and testing. The best model performance is yielded by distance measures which focus mainly on small perceptual distances and neglect outliers. (C) 2009 Acoustical Society of America. [DOI: 10.1121/1.3224721]