Efficient Personalized Speech Enhancement Through Self-Supervised Learning

Efficient Personalized Speech Enhancement Through Self-Supervised Learning
复制标题

基于自监督学习的高效个性化语音增强

DOI:
10.1109/jstsp.2022.3181782
复制
发表时间:
2021-04
影响因子:
7.5
通讯作者:
Aswin Sivaraman;Minje Kim
Aswin Sivaraman;Minje Kim
中科院分区:
工程技术1区
文献类型:
--
作者:
Aswin Sivaraman;Minje Kim

文献摘要

相似文献

这项工作提出了用于单声道说话者特定的自监督学习方法(即,个性化)语音增强模型。虽然通用模型必须广泛地解决许多扬声器,但个性化模型可以适应特定扬声器的声音,期望解决更窄的问题。因此,个性化除了降低计算复杂度之外,还可以实现更优化的性能。然而,幼稚的个性化方法可能不方便地需要来自目标用户的干净语音,例如,由于低于标准的记录条件。为此,我们将个性化设置为零拍摄任务,其中不使用目标说话者的干净语音,或者是少数拍摄学习任务,这是为了最大限度地减少用于迁移学习的干净语音的持续时间。在本文中,我们提出了自监督学习方法作为零次和少量个性化任务的解决方案。所提出的方法从未标记的数据(即,来自目标用户的野外噪声记录)而不是来自干净的源。我们研究了三种不同的自监督学习机制。我们建立了一个伪语音增强问题作为借口任务,它预训练模型来估计带噪语音,就好像它是干净的目标。对比学习和数据净化方法规范了伪增强问题的损失函数,克服了从未标记数据学习的局限性。我们通过将知名的ConvTasNet架构个性化到20个不同的目标扬声器来评估我们的方法。结果表明,基于自我监督的个性化提高了原始ConvTasNet的增强质量,具有更少的模型参数和更少的目标用户的干净数据。
This work presents self-supervised learning methods for monaural speaker-specific (i.e., personalized) speech enhancement models. While general-purpose models must broadly address many speakers, personalized models can adapt to a particular speaker's voice, expecting to solve a narrower problem. Hence, personalization can achieve more optimal performance in addition to reducing computational complexity. However, naive personalization methods can inconveniently require clean speech from the target user, e.g., due to subpar recording conditions. To this end, we pose personalization as either a zero-shot task, in which no clean speech of the target speaker is used, or a few-shot learning task, which is to minimize the duration of the clean speech used for transfer learning. With this paper, we propose self-supervised learning methods as a solution to both zero- and few-shot personalization tasks. The proposed methods learn the personalized speech features from unlabeled data (i.e., in-the-wild noisy recordings from the target user) rather than from the clean sources. We investigate three different self-supervised learning mechanisms. We set up a pseudo speech enhancement problem as a pretext task, which pretrains the models to estimate noisy speech as if it were the clean target. Contrastive learning and data purification methods regularize the loss function of the pseudo enhancement problem, overcoming the limitations of learning from unlabeled data. We assess our methods by personalizing the well-known ConvTasNet architecture to twenty different target speakers. The results show that self-supervision-based personalization improves the original ConvTasNet's enhancement quality with fewer model parameters and less clean data from the target user.