Factors influencing intelligibility of ideal binary-masked speech: Implications for noise reduction

Factors influencing intelligibility of ideal binary-masked speech: Implications for noise reduction
复制标题

DOI:
10.1121/1.2832617
复制
发表时间:
2008-03-01
影响因子:
2.4
通讯作者:
Loizou, Philipos C.
Loizou, Philipos C.
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Li, Ning;Loizou, Philipos C.

文献摘要

被引文献

相似文献

理想的二进制掩模的应用程序的听觉混合已被证明产生显着改善可懂度。该掩模通常应用于混合信号的时频(T-F)表示,并消除低于信噪比(SNR)阈值的信号部分,同时允许其他信号完整地通过。理想的二进制掩蔽语音的可懂度的影响因素还没有得到很好的理解,并在本研究中进行了检查。具体而言,本地SNR阈值,输入SNR水平,掩蔽类型,并在估计理想的掩模引入的错误的影响进行检查。与以前的研究一致,即使在-10 dB SNR下,所有测试的掩蔽物的二进制掩蔽刺激的可懂度也相当高。当掩蔽声主导的T-F单位被错误地标记为目标主导的T-F单位时,性能受到的影响最大。对于SNR阈值范围从-20到5 dB,性能稳定在接近100%的正确率。平台区域的存在表明,理想二进制掩模的图案最重要,而不是每个T-F单元的局部SNR。这种模式引导听者的注意力到目标的位置,并使他们能够在多人环境中有效地分离语音。(C)2008年,美国声学学会。
The application of the ideal binary mask to an auditory mixture has been shown to yield substantial improvements in intelligibility. This mask is commonly applied to the time-frequency (T-F) representation of a mixture signal and eliminates portions of a signal below a signal-to-noise-ratio (SNR) threshold while allowing others to pass through intact. The factors influencing intelligibility of ideal binary-masked speech are not well understood and are examined in the present study. Specifically, the effects of the local SNR threshold, input SNR level, masker type, and errors introduced in estimating the ideal mask are examined. Consistent with previous studies, intelligibility of binary-masked stimuli is quite high even at -10 dB SNR for all maskers tested. Performance was affected the most when the masker dominated T-F units were wrongly labeled as target-dominated T-F units, Performance plateaued near 100% correct for SNR thresholds ranging from -20 to 5 dB. The existence of the plateau region suggests that it is the pattern of the ideal binary mask that matters the most rather than the local SNR of each T-F unit. This pattern directs the listener's attention to where the target is and enables them to segregate speech effectively in multitalker environments. (C) 2008 Acoustical Society of America.