Quality assessment of speech enhancement systems by separation of enhanced speech, noise, and echo

Quality assessment of speech enhancement systems by separation of enhanced speech, noise, and echo
复制标题

通过分离增强语音、噪声和回声来评估语音增强系统的质量

DOI:
10.21437/interspeech.2007-307
复制
发表时间:
2007
期刊:
Autom.
影响因子:
--
通讯作者:
Suhadi Suhadi
Suhadi Suhadi
中科院分区:
--
文献类型:
--
作者:
T. Fingscheidt;Suhadi Suhadi

文献摘要

被引文献

相似文献

摘要语音增强系统的质量评估是一项重要的任务,特别是当(残余)噪声和回声信号分量出现时。我们提出了一个信号分离schemethat允许在黑盒测试场景中的未知语音增强系统的详细分析。我们的方法分离的语音,(残留)噪声,(残留)回声成分的语音增强系统在发送方向(上行链路方向)。这使得可以独立地判断语音退化和噪声和回声衰减/退化。虽然最先进的测试总是试图判断发送方向信号的混合,我们的新方案允许在更短的时间内更可靠的分析。这将是非常有用的测试免提设备在实践中,以及测试语音增强算法的研究和发展。指标术语:客观信号质量评估,非盲信号分离,语音增强,免提1。在科学上,评估语音增强算法的一种舒适的方法是将近端语音和噪声数字化地添加到回声信号中,从而构建麦克风信号。在语音增强(免提)系统的上行链路处理期间,然后记录对有噪声的麦克风信号的操作干扰,并且稍后单独地应用于麦克风信号的语音、回声和噪声分量(参见,例如,[1、2、3])。这假定线性处理,如可以在例如频域降噪中发现的,其中再次应用于频谱幅度。这种方法的优点在于可以获得三个独立的信号:填充语音分量、填充回声分量和填充噪声分量,它们分别表示(轻微)失真的近端说话者的语音信号、抑制回声信号和残余噪声信号。专注于降噪,例如,语音失真、噪声衰减和噪声失真等方面可以进行舒适的测量或听觉评估伊萨
Abstract Quality assessment of speech enhancement systems is a non-trivial task, especially when (residual) noise and echo signalcomponents occur. We present a signal separation schemethat allows for a detailed analysis of unknown speech enhance-ment systems in a black box test scenario. Our approach sep-arates the speech, (residual) noise, and (residual) echo compo-nent of the speech enhancement system in the sending direc-tion (uplink direction). This makes it possible to independentlyjudge the speech degradation and the noise and echo attenua-tion/degradation. While state of the art tests always try to judgethe sending direction signal mixture, our new scheme allows amore reliable analysis in shorter time. It will be very usefulfor testing hands-free devices in practice as well as for testingspeech enhancement algorithms in research and development. Index Terms : objective signal quality assessment, non-blindsignal separation, speech enhancement, hands-free 1. Introduction In science, a comfortable way to evaluate speech enhancementalgorithms is to digitally add near-end speech and noise to theecho signal and thereby construct the microphone signal. Dur-ing the uplink processing of the speech enhancement (hands-free) system the operational influence on the noisy microphonesignal is then to be logged, and later applied individually to thespeech, echo, and noise components of the microphone signal(see, e.g., [1, 2, 3]). This presumes linear processing, as canbe found e.g. in frequency domain noise reduction, where again is applied to the spectral amplitudes. The strength of suchmethod is that one achieves three separate signals: The filteredspeechcomponent, thefilteredechocomponent, andthefilterednoise component, which represent the (slightly) distorted near-end talker’s speech signal, the suppressed echo signal, and theresidual noisesignal, respectively. Focusingonnoisereduction,e.g., aspects such as speech distortion, noise attenuation, andnoisedistortioncan thencomfortably bemeasured or auditivelyassessed.Thishowever isa