Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation

Combining Spectral and Self-Supervised Features for Low Resource Speech Recognition and Translation
复制标题

DOI:
10.48550/arxiv.2204.02470
复制
发表时间:
2022-04
期刊:
--
影响因子:
--
通讯作者:
Dan Berrebbi;Jiatong Shi;Brian Yan;Osbel López-Francisco;Jonathan D. Amith;Shinji Watanabe
Dan Berrebbi;Jiatong Shi;Brian Yan;Osbel López-Francisco;Jonathan D. Amith;Shinji Watanabe
中科院分区:
其他
文献类型:
--
作者:
Dan Berrebbi;Jiatong Shi;Brian Yan;Osbel López-Francisco;Jonathan D. Amith;Shinji Watanabe

文献摘要

被引文献

相似文献

自监督学习(SSL)模型已成功应用于各种基于深度学习的语音任务,特别是那些数据量有限的任务。然而,SSL 表示的质量在很大程度上取决于 SSL 训练域和目标数据域之间的相关性。相反,谱特征(SF)提取器(例如对数梅尔滤波器组)是手工制作的不可学习组件,并且对于域移位可能更加鲁棒。目前的工作检验了这样的假设:将不可学习的 SF 提取器与 SSL 模型相结合是处理低资源语音任务的有效方法。我们提出了一个可学习和可解释的框架来结合 SF 和 SSL 表示。所提出的框架在三个低资源数据集上的自动语音识别 (ASR) 和语音翻译 (ST) 任务上显着优于基线模型和 SSL 模型。我们还设计了基于专家的组合模型。最后一个模型表明,在 SSL 训练集和目标语言数据之间域不匹配的情况下,SSL 模型相对于传统 SF 提取器的相对贡献非常小。
Self-Supervised Learning (SSL) models have been successfully applied in various deep learning-based speech tasks, particularly those with a limited amount of data. However, the quality of SSL representations depends highly on the relatedness between the SSL training domain(s) and the target data domain. On the contrary, spectral feature (SF) extractors such as log Mel-filterbanks are hand-crafted non-learnable components, and could be more robust to domain shifts. The present work examines the assumption that combining non-learnable SF extractors to SSL models is an effective approach to low resource speech tasks. We propose a learnable and interpretable framework to combine SF and SSL representations. The proposed framework outperforms significantly both baseline and SSL models on Automatic Speech Recognition (ASR) and Speech Translation (ST) tasks on three low resource datasets. We additionally design a mixture of experts based combination model. This last model reveals that the relative contribution of SSL models over conventional SF extractors is very small in case of domain mismatch between SSL training set and the target language data.