Frequency domain binaural model based on interaural phase and level differences

Frequency domain binaural model based on interaural phase and level differences
复制标题

基于耳间相位和电平差的频域双耳模型

DOI:
10.1250/ast.24.172
复制
发表时间:
2003
影响因子:
0.7
通讯作者:
M. Ebata
M. Ebata
中科院分区:
--
文献类型:
--
作者:
H. Nakashima;Y. Chisaki;T. Usagawa;M. Ebata

文献摘要

被引文献

相似文献

我们可以在嘈杂的环境中与他人交流。这种现象被称为“鸡尾酒会效应”,是最重要的双耳功能之一。本文提出了一种频域双耳模型,该模型基于两耳间的相位和电平差,起到双耳功能的作用。该模型不仅作为语音识别系统的前端,而且作为语音增强器进行评估。根据该评估,当目标信号和噪声的到达方向相差10°时,在任何情况下与先前的时域双耳模型(TDBM)相比,识别率都有所提高。此外,当信噪比(SNR)高于约5 dB时,识别率超过90%。另一方面,SNR和相干性的频域双耳模型,这是获得的语音增强器的评估,显示优于TDBM的上级的结果。
We can communicate with others in a noisy environment. This phenomenon is known as a “Cocktail Party Effect” and is one of the most important binaural functions. This paper addresses a frequency domain binaural model that plays the role of a binaural function based on an interaural phase and level difference. The proposed model is evaluated not only as a front-end of the speech recognition system, but also as a speech enhancer. According to the evaluation, when the direction of arrival of the target signal and noise differs by 10°, recognition rates improve in comparison with the previous time domain binaural model (TDBM) in any cases. Furthermore, recognition rates show more than 90% when the signal to noise ratio (SNR) is higher than approximately 5 dB. On the other hand, SNR and coherence of the frequency domain binaural model, which is obtained for an evaluation of the speech enhancer, show superior results over the TDBM.