Data Augmenting Contrastive Learning of Speech Representations in the Time Domain

Data Augmenting Contrastive Learning of Speech Representations in the Time Domain
复制标题

DOI:
10.1109/slt48900.2021.9383605
复制
发表时间:
2020-07
期刊:
2021 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
E. Kharitonov;M. Rivière;Gabriel Synnaeve;Lior Wolf;Pierre-Emmanuel Mazar'e;Matthijs Douze;Emmanuel Dupoux
E. Kharitonov;M. Rivière;Gabriel Synnaeve;Lior Wolf;Pierre-Emmanuel Mazar'e;Matthijs Douze;Emmanuel Dupoux
中科院分区:
其他
文献类型:
--
作者:
E. Kharitonov;M. Rivière;Gabriel Synnaeve;Lior Wolf;Pierre-Emmanuel Mazar'e;Matthijs Douze;Emmanuel Dupoux

文献摘要

被引文献

相似文献

对比预测编码(CPC)是一种基于从过去的语音片段中预测未来的语音片段的算法,是语音信号表示学习的一种强大算法。然而,与其他方法相比,它在无监督评估基准上仍然表现不佳。在这里,我们介绍了WavAugment,这是一个时域数据增强库,我们根据CPC(原始波形输入,对比损耗,过去与未来结构)的特殊性进行了调整和优化。我们发现,仅对执行CPC预测的部分应用增强比将其应用于提取对比损失样本(正负)的未来部分产生更好的结果。在librisspeech的无监督指标上选择了音高修改、加性噪声和混响的最佳组合后(相对于ABX分数的增益为18-22%),我们将这种组合不加任何更改地应用于零资源语音基准2017中的三个新数据集,并使用域外训练数据击败了最先进的数据集。最后,我们表明,数据增强预训练特征改善了lib -light半监督设置(标记数据10分钟、1小时或10小时)下的下游电话识别任务,相对降低了15%的PER。
Contrastive Predictive Coding (CPC), based on predicting future segments of speech from past segments is emerging as a powerful algorithm for representation learning of speech signal. However, it still under-performs compared to other methods on unsupervised evaluation benchmarks. Here, we intro-duce WavAugment, a time-domain data augmentation library which we adapt and optimize for the specificities of CPC (raw waveform input, contrastive loss, past versus future structure). We find that applying augmentation only to the segments from which the CPC prediction is performed yields better results than applying it also to future segments from which the samples (both positive and negative) of the contrastive loss are drawn. After selecting the best combination of pitch modification, additive noise and reverberation on unsupervised metrics on LibriSpeech (with a gain of 18-22% relative on the ABX score), we apply this combination without any change to three new datasets in the Zero Resource Speech Benchmark 2017 and beat the state-of-the-art using out-of-domain training data. Finally, we show that the data-augmented pretrained features improve a downstream phone recognition task in the Libri-light semi-supervised setting (10 min, 1 h or 10 h of labelled data) reducing the PER by 15% relative.