Voice activity detection with noise reduction and long-term spectral divergence estimation

Voice activity detection with noise reduction and long-term spectral divergence estimation
复制标题

具有降噪和长期频谱散度估计的语音活动检测

DOI:
10.1109/icassp.2004.1326452
复制
发表时间:
2004
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
A. Rubio
A. Rubio
中科院分区:
--
文献类型:
--
作者:
J. Ramírez;J. C. Segura;M. C. Benítez;Á. D. L. Torre;A. Rubio

文献摘要

被引文献

相似文献

本文主要研究了一种改进的基于长时信号处理和最大谱跟踪的语音端点检测算法。这种方法的益处已经在先前的工作中进行了分析(Ramirez,J.等人,Proc. EUROSPEECH 2003,p.3041-4,2003),在噪声环境中的语音/非语音可辨别性和语音识别性能方面具有明显的改进。现在考虑两个明确的方面。第一个,这提高了在低噪声条件下的VAD的性能,考虑了自适应长度的帧窗口来跟踪长期频谱分量。第二种方法通过在长期谱跟踪之前使用降噪阶段来减少高噪声环境中的误分类错误。实验结果表明,不同的语音/停顿的歧视和语音识别性能的VAD方法有明显的改善。特别是,当建议的VAD取代了分布式语音识别(DSR)的ETSI高级前端(AFE)的VAD时,识别率有所提高。
The paper mainly focusses on an improved voice activity detection algorithm employing long-term signal processing and maximum spectral component tracking. The benefits of this approach have been analyzed in a previous work (Ramirez, J. et al., Proc. EUROSPEECH 2003, p.3041-4, 2003) with clear improvements in speech/non-speech discriminability and speech recognition performance in noisy environments. Two clear aspects are now considered. The first one, which improves the performance of the VAD in low noise conditions, considers an adaptive length frame window to track the long-term spectral components. The second one reduces misclassification errors in highly noisy environments by using a noise reduction stage before the long-term spectral tracking. Experimental results show clear improvements over different VAD methods in speech/pause discrimination and speech recognition performance. Particularly, improvements in recognition rate were reported when the proposed VAD replaced the VADs of the ETSI advanced front-end (AFE) for distributed speech recognition (DSR).