Voice Activity Detection in the Wild: A Data-Driven Approach Using Teacher-Student Training

Voice Activity Detection in the Wild: A Data-Driven Approach Using Teacher-Student Training
复制标题

野外语音活动检测:使用师生培训的数据驱动方法

DOI:
10.1109/taslp.2021.3073596
复制
发表时间:
2021-05
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Kai Yu
Kai Yu
中科院分区:
其他
文献类型:
--
作者:
Heinrich Dinkel;Shuai Wang;Xuenan Xu;Mengyue Wu;Kai Yu

文献摘要

参考文献

相似文献

语音活动检测是自动语音识别(ASR)等语音相关任务的重要预处理组件。传统的监督式VAD系统通过使用例如隐马尔可夫模型这些ASR模型通常在干净和完全转录的数据上训练,限制VAD系统在干净或合成噪声数据集上训练。因此,有监督的VAD系统的一个主要挑战是它们对有噪声的真实世界数据的泛化。这项工作提出了一种数据驱动的教师-学生的VAD方法,它利用大量的和不受约束的音频数据进行训练。与以前的方法不同,在教师培训期间只需要弱标签,从而可以利用任何真实世界的潜在噪声数据集。我们的方法首先使用剪辑级监督在源数据集(Audioset)上训练教师模型。训练后,教师在未标记的目标数据集上为学生模型提供帧级指导。研究了在中型到大型数据集上训练的大量学生模型(Audioset,Voxceleb,NIST SRE)。然后,我们的方法分别在干净的,人为噪声和真实世界的数据进行评估。我们观察到显着的性能增益在人为噪声和现实世界的情况下。最后,我们将我们的方法与其他无监督和有监督的VAD方法进行了比较,证明了我们方法的优越性。
Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a Hidden Markov model. These ASR models are commonly trained on clean and fully transcribed data, limiting VAD systems to be trained on clean or synthetically noised datasets. Therefore, a major challenge for supervised VAD systems is their generalization towards noisy, real-world data. This work proposes a data-driven teacher-student approach for VAD, which utilizes vast and unconstrained audio data for training. Unlike previous approaches, only weak labels during teacher training are required, enabling the utilization of any real-world, potentially noisy dataset. Our approach firstly trains a teacher model on a source dataset (Audioset) using clip-level supervision. After training, the teacher provides frame-level guidance to a student model on an unlabeled, target dataset. A multitude of student models trained on mid- to large-sized datasets are investigated (Audioset, Voxceleb, NIST SRE). Our approach is then respectively evaluated on clean, artificially noised, and real-world data. We observe significant performance gains in artificially noised and real-world scenarios. Lastly, we compare our approach against other unsupervised and supervised VAD methods, demonstrating our method's superiority.
DOI: 10.1016/j.csl.2019.06.005
发表时间: 2020-01-01
影响因子: 4.3
作者:
Tan, Zheng-Hua;Sarkar, Achintya Kr;Dehak, Najim
通讯作者: Dehak, Najim
DOI: 10.1109/icassp.2011.5947431
发表时间: 2011-05
期刊: 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
J. A. Morales-Cordovilla;Ning Ma;V. Sánchez;J. L. Carmona;A. Peinado;J. Barker
通讯作者: J. A. Morales-Cordovilla;Ning Ma;V. Sánchez;J. L. Carmona;A. Peinado;J. Barker
DOI: 10.1109/icassp.2018.8461975
发表时间: 2017-10
期刊: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Yong Xu;Qiuqiang Kong;Wenwu Wang;Mark D. Plumbley
通讯作者: Yong Xu;Qiuqiang Kong;Wenwu Wang;Mark D. Plumbley
DOI: 10.1109/tasl.2011.2125953
发表时间: 2011-11
期刊: IEEE Transactions on Audio, Speech, and Language Processing
影响因子: --
作者:
D. Ying;Yonghong Yan;J. Dang;F. Soong
通讯作者: D. Ying;Yonghong Yan;J. Dang;F. Soong
DOI: 10.21437/interspeech.2020-0995
发表时间: 2020-03
期刊: --
影响因子: --
作者:
Yefei Chen;Heinrich Dinkel;Mengyue Wu;Kai Yu
通讯作者: Yefei Chen;Heinrich Dinkel;Mengyue Wu;Kai Yu