Test-Time Adaptation Toward Personalized Speech Enhancement: Zero-Shot Learning with Knowledge Distillation

Test-Time Adaptation Toward Personalized Speech Enhancement: Zero-Shot Learning with Knowledge Distillation
复制标题

DOI:
10.1109/waspaa52581.2021.9632771
复制
发表时间:
2021-05
期刊:
2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
影响因子:
--
通讯作者:
Sunwoo Kim;Minje Kim
Sunwoo Kim;Minje Kim
中科院分区:
其他
文献类型:
--
作者:
Sunwoo Kim;Minje Kim

文献摘要

被引文献

相似文献

在最终用户设备的实际语音增强设置中,我们经常只遇到少数扬声器和噪声类型,这些扬声器和噪声类型往往会在特定的声学环境中重复出现。我们提出了一种新的个性化语音增强方法,以适应一个紧凑的去噪模型的测试时间特异性。我们在这个测试时间适应的目标是利用没有干净的语音目标的测试发言人,从而满足零拍摄学习的要求。为了弥补干净语音的不足,我们采用了知识蒸馏框架:我们从一个过大的教师模型中提取更高级的去噪结果,并将其用作伪目标来训练小学生模型。这种零触发学习过程规避了收集用户的干净语音的过程,由于隐私问题和记录干净语音的技术困难,用户不愿意遵守该过程。在不同测试时间条件下的实验表明,所提出的个性化方法可以显着提高紧凑模型在测试时间内的性能。此外,由于个性化模型优于较大的非个性化基线模型,我们声称个性化实现了模型压缩,而不会损失去噪性能。正如预期的那样,学生模型的表现低于最先进的教师模型。
In realistic speech enhancement settings for end-user devices, we often encounter only a few speakers and noise types that tend to reoccur in the specific acoustic environment. We propose a novel personalized speech enhancement method to adapt a compact denoising model to the test-time specificity. Our goal in this test-time adaptation is to utilize no clean speech target of the test speaker, thus fulfilling the requirement for zero-shot learning. To complement the lack of clean speech, we employ the knowledge distillation framework: we distill the more advanced denoising results from an overly large teacher model, and use them as the pseudo target to train the small student model. This zero-shot learning procedure circumvents the process of collecting users' clean speech, a process that users are reluctant to comply due to privacy concerns and technical difficulty of recording clean voice. Experiments on various test-time conditions show that the proposed personalization method can significantly improve the compact models' performance during the test time. Furthermore, since the personalized models outperform larger non-personalized baseline models, we claim that personalization achieves model compression with no loss of denoising performance. As expected, the student models underperform the state-of-the-art teacher models.