Voice Activity Detection in the Wild: A Data-Driven Approach Using Teacher-Student Training
Voice Activity Detection in the Wild: A Data-Driven Approach Using Teacher-Student Training
复制标题
野外语音活动检测:使用师生培训的数据驱动方法
DOI:
10.1109/taslp.2021.3073596
复制
发表时间:
2021-05
期刊:
影响因子:
--
通讯作者:
Kai Yu
中科院分区:
文献类型:
--
作者:
Heinrich Dinkel;Shuai Wang;Xuenan Xu;Mengyue Wu;Kai Yu
Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a Hidden Markov model. These ASR models are commonly trained on clean and fully transcribed data, limiting VAD systems to be trained on clean or synthetically noised datasets. Therefore, a major challenge for supervised VAD systems is their generalization towards noisy, real-world data. This work proposes a data-driven teacher-student approach for VAD, which utilizes vast and unconstrained audio data for training. Unlike previous approaches, only weak labels during teacher training are required, enabling the utilization of any real-world, potentially noisy dataset. Our approach firstly trains a teacher model on a source dataset (Audioset) using clip-level supervision. After training, the teacher provides frame-level guidance to a student model on an unlabeled, target dataset. A multitude of student models trained on mid- to large-sized datasets are investigated (Audioset, Voxceleb, NIST SRE). Our approach is then respectively evaluated on clean, artificially noised, and real-world data. We observe significant performance gains in artificially noised and real-world scenarios. Lastly, we compare our approach against other unsupervised and supervised VAD methods, demonstrating our method's superiority.
登录
查看更多内容
影响因子:
4.3
作者:
Tan, Zheng-Hua;Sarkar, Achintya Kr;Dehak, Najim
通讯作者:
Dehak, Najim
DOI:
10.1109/icassp.2011.5947431
发表时间:
2011-05
期刊:
2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
J. A. Morales-Cordovilla;Ning Ma;V. Sánchez;J. L. Carmona;A. Peinado;J. Barker
通讯作者:
J. A. Morales-Cordovilla;Ning Ma;V. Sánchez;J. L. Carmona;A. Peinado;J. Barker
DOI:
10.1109/icassp.2018.8461975
发表时间:
2017-10
期刊:
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
Yong Xu;Qiuqiang Kong;Wenwu Wang;Mark D. Plumbley
通讯作者:
Yong Xu;Qiuqiang Kong;Wenwu Wang;Mark D. Plumbley
DOI:
10.1109/tasl.2011.2125953
发表时间:
2011-11
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
D. Ying;Yonghong Yan;J. Dang;F. Soong
通讯作者:
D. Ying;Yonghong Yan;J. Dang;F. Soong
DOI:
10.21437/interspeech.2020-0995
发表时间:
2020-03
期刊:
--
影响因子:
--
作者:
Yefei Chen;Heinrich Dinkel;Mengyue Wu;Kai Yu
通讯作者:
Yefei Chen;Heinrich Dinkel;Mengyue Wu;Kai Yu