Close/distant talker discrimination based on kurtosis of linear prediction residual signals

Close/distant talker discrimination based on kurtosis of linear prediction residual signals
复制标题

DOI:
10.1109/icassp.2014.6854015
复制
发表时间:
2014-05
期刊:
2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Kohei Hayashida;M. Nakayama;T. Nishiura;Y. Yamashita;T. Horiuchi;T. Kato
Kohei Hayashida;M. Nakayama;T. Nishiura;Y. Yamashita;T. Horiuchi;T. Kato
中科院分区:
其他
文献类型:
--
作者:
Kohei Hayashida;M. Nakayama;T. Nishiura;Y. Yamashita;T. Horiuchi;T. Kato

文献摘要

相似文献

为了实现有用的应用,如语音接口和电话会议系统,期望/不希望的语音歧视与语音/非语音歧视一样重要。传统的语音活动检测(VAD)方法利用声源的方向信息来区分需要的语音和不需要的语音。然而,这些方法必须利用多个麦克风来估计声源的方向。在这里,我们提出了一种新的方法来区分需要和不需要的语音在一个单一的麦克风。我们假设期望的说话者离麦克风很近,提出的方法可以根据线性预测(LP)残差信号的峰度从观测信号中区分出近/远的说话语音。实验结果表明,在普通混响环境下,该方法能够在10%的等误差率(EER)内识别出近距离说话和远距离说话的语音,且处理时间短。
Desired/undesired speech discrimination is as important as speech/non-speech discrimination to achieve useful applications such as speech interfaces and teleconferencing systems. Conventional methods of voice activity detection (VAD) utilize the directional information of sound sources to distinguish desired from undesired speech. However, these methods have to utilize multiple microphones to estimate the directions of sound sources. Here, we propose a new method to discriminate desired from undesired speech with a single microphone. We assumed that the desired talkers would be close to the microphone, and the proposed method could distinguish close/distant-talking speech from observed signals based on the kurtosis of the linear prediction (LP) residual signals. The experimental results revealed that the proposed method could distinguish close-talking speech from distant-talking speech within a 10% equal error rate (EER) in ordinary reverberant environments with less processing time.