Voice Activity Detection Based on an Unsupervised Learning Framework

Voice Activity Detection Based on an Unsupervised Learning Framework
复制标题

DOI:
10.1109/tasl.2011.2125953
复制
发表时间:
2011-11
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
D. Ying;Yonghong Yan;J. Dang;F. Soong
D. Ying;Yonghong Yan;J. Dang;F. Soong
中科院分区:
其他
文献类型:
--
作者:
D. Ying;Yonghong Yan;J. Dang;F. Soong

文献摘要

被引文献

相似文献

如何建立语音/非语音识别模型是语音活动检测器(VAD)的关键。半监督学习是传统VAD中最流行的模型构建方法。在这封信中,我们提出了一个无监督的学习框架来构建VAD的统计模型。这个框架是通过一个连续的高斯混合模型来实现的。它包括初始化过程和更新过程。在每个子带,GMM首先使用EM算法初始化,然后逐帧顺序更新。从GMM中,在每个子带处导出用于辨别的自我调节阈值。为了保证GMM的可靠性,本文引入了一些约束条件。由于无监督学习的原因,所提出的VAD不依赖于一个假设,即话语的前几帧是非语音,这是广泛使用的大多数VAD。此外,在时间-频率域中的语音存在概率是该VAD的副产品。我们测试了语音从TIMIT数据库和噪声从NOISEX-92数据库。与ITU G.729B、GSM AMR和典型的半监督VAD相比,评估有效地显示了其有前途的性能。
How to construct models for speech/nonspeech discrimination is a crucial point for voice activity detectors (VADs). Semi-supervised learning is the most popular way for model construction in conventional VADs. In this correspondence, we propose an unsupervised learning framework to construct statistical models for VAD. This framework is realized by a sequential Gaussian mixture model. It comprises an initialization process and an updating process. At each subband, the GMM is firstly initialized using EM algorithm, and then sequentially updated frame by frame. From the GMM, a self-regulatory threshold for discrimination is derived at each subband. Some constraints are introduced to this GMM for the sake of reliability. For the reason of unsupervised learning, the proposed VAD does not rely on an assumption that the first several frames of an utterance are nonspeech, which is widely used in most VADs. Moreover, the speech presence probability in the time-frequency domain is a byproduct of this VAD. We tested it on speech from TIMIT database and noise from NOISEX-92 database. The evaluations effectively showed its promising performance in comparison with VADs such as ITU G.729B, GSM AMR, and a typical semi-supervised VAD.