Comparative Layer-Wise Analysis of Self-Supervised Speech Models

Comparative Layer-Wise Analysis of Self-Supervised Speech Models
复制标题

DOI:
10.1109/icassp49357.2023.10096149
复制
发表时间:
2022-11
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Ankita Pasad;Bowen Shi;Karen Livescu
Ankita Pasad;Bowen Shi;Karen Livescu
中科院分区:
其他
文献类型:
--
作者:
Ankita Pasad;Bowen Shi;Karen Livescu

文献摘要

被引文献

相似文献

在过去的几年里,已经提出了许多自监督语音模型,其预训练目标,输入模态和预训练数据各不相同。尽管在下游任务上取得了令人印象深刻的成功,但我们对模型编码的属性以及模型之间的差异仍然了解有限。在这项工作中,我们研究了各种最近的模型的中间表示。具体来说,我们测量声学,语音和单词级属性编码在各个层中,使用一个轻量级的分析工具,基于典型相关分析(CCA)。我们发现,这些属性在不同的层之间根据模型的不同而变化,这些变化与预训练目标的选择有关。我们进一步研究我们的分析,通过比较性能的趋势与语音识别和口语理解任务的下游任务的效用。我们发现,CCA趋势为下游任务选择感兴趣的层提供了可靠的指导,并且单层性能通常与使用所有层相匹配或有所改善,这意味着更有效地使用预训练模型。1
Many self-supervised speech models, varying in their pre-training objective, input modality, and pre-training data, have been proposed in the last few years. Despite impressive successes on downstream tasks, we still have a limited understanding of the properties encoded by the models and the differences across models. In this work, we examine the intermediate representations for a variety of recent models. Specifically, we measure acoustic, phonetic, and word-level properties encoded in individual layers, using a lightweight analysis tool based on canonical correlation analysis (CCA). We find that these properties evolve across layers differently depending on the model, and the variations relate to the choice of pre-training objective. We further investigate the utility of our analyses for downstream tasks by comparing the property trends with performance on speech recognition and spoken language understanding tasks. We discover that CCA trends provide reliable guidance to choose layers of interest for downstream tasks and that single-layer performance often matches or improves upon using all layers, suggesting implications for more efficient use of pre-trained models. 1