SNR Loss: A new objective measure for predicting speech intelligibility of noise-suppressed speech.

SNR Loss: A new objective measure for predicting speech intelligibility of noise-suppressed speech.
复制标题

DOI:
10.1016/j.specom.2010.10.005
复制
发表时间:
2011-03-01
影响因子:
3.2
通讯作者:
Loizou, Philipos C.
Loizou, Philipos C.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ma, Jianfen;Loizou, Philipos C.

文献摘要

参考文献

被引文献

相似文献

大多数现有的可懂度度量方法没有考虑到处理后的语音中存在的失真,例如由语音增强算法引入的失真。在本研究中,我们提出了三种新的客观度量方法,可用于预测在噪声条件下经过处理(例如通过增强算法)的语音的可懂度。这三种度量方法都使用了纯净信号和噪声抑制信号的临界频带频谱表示,并且基于在受损信号经过语音增强算法后每个临界频带中产生的信噪比损失的测量。所提出的度量方法具有灵活性,因为它们可以对增强算法引入的两种频谱失真(即频谱衰减和频谱放大失真)赋予不同的权重。通过听力正常的受试者在72种噪声条件下获得的可懂度分数对所提出的度量方法进行了评估,这些噪声条件涉及被四种不同的掩蔽声(汽车、嘈杂人声、火车和街道干扰)干扰的噪声抑制语音(辅音和句子)。使用一种仅包含元音/辅音过渡和弱辅音信息的信噪比损失度量方法的变体,与句子识别分数获得了最高的相关性(r = -0.85)。对于所有噪声类型都保持了高相关性,在街道噪声条件下达到了最大相关性(r = -0.88)。
Most of the existing intelligibility measures do not account for the distortions present in processed speech, such as those introduced by speech-enhancement algorithms. In the present study, we propose three new objective measures that can be used for prediction of intelligibility of processed (e.g., via an enhancement algorithm) speech in noisy conditions. All three measures use a critical-band spectral representation of the clean and noise-suppressed signals and are based on the measurement of the SNR loss incurred in each critical band after the corrupted signal goes through a speech enhancement algorithm. The proposed measures are flexible in that they can provide different weights to the two types of spectral distortions introduced by enhancement algorithms, namely spectral attenuation and spectral amplification distortions. The proposed measures were evaluated with intelligibility scores obtained by normal-hearing listeners in 72 noisy conditions involving noise-suppressed speech (consonants and sentences) corrupted by four different maskers (car, babble, train and street interferences). Highest correlation (r=−0.85) with sentence recognition scores was obtained using a variant of the SNR loss measure that only included vowel/consonant transitions and weak consonant information. High correlation was maintained for all noise types, with a maximum correlation (r=−0.88) achieved in street noise conditions.
DOI: 10.1109/tsa.2003.814458
发表时间: 2003-07-01
期刊: IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
影响因子: --
作者:
Hu, Y;Loizou, PC
通讯作者: Loizou, PC
DOI: 10.1109/tasl.2008.919072
发表时间: 2008-05-01
影响因子: --
作者:
Benesty, Jacob;Chen, Jingdong;Huang, Yiteng (Arden)
通讯作者: Huang, Yiteng (Arden)
DOI: 10.1121/1.392224
发表时间: 1985-01-01
影响因子: 2.4
作者:
HOUTGAST, T;STEENEKEN, HJM
通讯作者: STEENEKEN, HJM
DOI: 10.1109/tsa.2005.860851
发表时间: 2006-07-01
影响因子: --
作者:
Chen, Jingdong;Benesty, Jacob;Doclo, Simon
通讯作者: Doclo, Simon
DOI: 10.1109/89.326615
发表时间: 1994-10-01
期刊: IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
影响因子: --
作者:
Allen, Jont B.
通讯作者: Allen, Jont B.