Modelling speaker intelligibility in noise

Modelling speaker intelligibility in noise
复制标题

DOI:
10.1016/j.specom.2006.11.003
复制
发表时间:
2007-05-01
影响因子:
3.2
通讯作者:
Cooke, Martin
Cooke, Martin
中科院分区:
计算机科学3区
文献类型:
--
作者:
Barker, Jon;Cooke, Martin

文献摘要

被引文献

相似文献

这项研究将听众在多说话者噪声中语音任务中的表现与受自动语音识别技术启发的模型的表现进行了比较。听众在一系列信噪比的语音形状噪声中识别出简单的 6 个单词句子中的三个关键词。句子材料由 18 名男性或 16 名女性发言者提供。对许多声学参数(声道长度、平均基频和语速)的跨说话者分析发现,没有一个能够始终如一地很好地预测相对清晰度。能量掩蔽程度的简单测量可以很好地预测女性语音清晰度,尤其是在高噪声条件下,但无法解释男性群体中说话者之间的差异。一个将能量掩蔽模拟与说话者相关统计模型相结合的一瞥模型,产生了适合所有说话者汇集的行为数据的识别分数。使用一组与说话者无关、与噪声水平无关的参数,该模型不仅能够在很大程度上预测各个说话者的可懂度,而且还可以解释字母关键字的大部分标记式可懂度。在高噪音条件下,配合效果特别好。 (c) 2006 Elsevier B.V. 保留所有权利。
This study compared listeners' performance on a multispeaker speech-in-noise task with that of a model inspired by automatic speech recognition techniques. Listeners identified three keywords in simple 6-word sentences presented in speech-shaped noise at a range of signal-to-noise ratios. Sentence material was provided by 18 male or 16 female speakers. An across-speaker analysis of a number of acoustic parameters (vocal tract length, mean fundamental frequency and speaking rate) found none to be consistently good predictors of relative intelligibility. A simple measure of degree of energetic masking was a good predictor of female speech intelligibility, especially in high noise conditions, but failed to account for interspeaker differences for the male group. A glimpsing model, which combined a simulation of energetic masking with speaker-dependent statistical models, produced recognition scores which were fitted to the behavioural data pooled across all speakers. Using a single set of speaker-independent, noise-level-independent parameters, the model was able to predict not only the intelligibility of individual speakers to a remarkable degree, but could also account for most of the token-wise intelligibilities of the letter keywords. The fit was particularly good in high noise conditions. (c) 2006 Elsevier B.V. All rights reserved.