Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification

Personalized Speech Enhancement through Self-Supervised Data Augmentation and Purification
复制标题

DOI:
10.21437/interspeech.2021-1868
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Aswin Sivaraman;Sunwoo Kim;Minje Kim
Aswin Sivaraman;Sunwoo Kim;Minje Kim
中科院分区:
其他
文献类型:
--
作者:
Aswin Sivaraman;Sunwoo Kim;Minje Kim

文献摘要

相似文献

由于隐私限制和目标用户对无噪声语音的有限访问,训练个性化语音增强模型天生就是一个无机会学习问题。如果测试时间用户有大量未标记的有噪语音,则可以使用自监督学习来训练个性化语音增强模型。一种直接的个性化建模方法是使用目标说话人的嘈杂录音作为伪源。然后,伪去噪模型学习去除注入的训练噪声并恢复伪源。然而,这种方法是不稳定的,因为它取决于伪源的质量,这可能太过噪声。作为补救措施,我们提出了一种通过数据净化来改进自我监督的方法。我们首先训练一个信噪比预测器模型来估计伪源的逐帧信噪比。然后,预测器的估计被转换成权重,该权重调整伪源对训练个性化模型的逐帧贡献。我们的实验表明,在个性化语音增强的背景下,所提出的数据提纯步骤提高了特定说话人噪声数据的可用性。在不依赖任何干净的语音记录或说话人嵌入的情况下,我们的方法可能被视为隐私保护。
Training personalized speech enhancement models is innately a no-shot learning problem due to privacy constraints and limited access to noise-free speech from the target user. If there is an abundance of unlabeled noisy speech from the test-time user, a personalized speech enhancement model can be trained using self-supervised learning. One straightforward approach to model personalization is to use the target speaker's noisy recordings as pseudo-sources. Then, a pseudo denoising model learns to remove injected training noises and recover the pseudo-sources. However, this approach is volatile as it depends on the quality of the pseudo-sources, which may be too noisy. As a remedy, we propose an improvement to the self-supervised approach through data purification. We first train an SNR predictor model to estimate the frame-by-frame SNR of the pseudo-sources. Then, the predictor's estimates are converted into weights which adjust the frame-by-frame contribution of the pseudo-sources towards training the personalized model. We empirically show that the proposed data purification step improves the usability of the speaker-specific noisy data in the context of personalized speech enhancement. Without relying on any clean speech recordings or speaker embeddings, our approach may be seen as privacy-preserving.