Layer-Wise Analysis of a Self-Supervised Speech Representation Model

Layer-Wise Analysis of a Self-Supervised Speech Representation Model
复制标题

DOI:
10.1109/asru51503.2021.9688093
复制
发表时间:
2021-07
期刊:
2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
影响因子:
--
通讯作者:
Ankita Pasad;Ju-Chieh Chou;Karen Livescu
Ankita Pasad;Ju-Chieh Chou;Karen Livescu
中科院分区:
其他
文献类型:
--
作者:
Ankita Pasad;Ju-Chieh Chou;Karen Livescu

文献摘要

被引文献

相似文献

最近提出的自监督学习方法已经成功地用于预训练语音表示模型。这些习得的表征的效用已经被经验地观察到了,但关于预先训练的表征本身所编码的信息的类型或范围的研究并不多。开发这样的洞察力可以帮助理解这些模型的功能和限制,并使研究社区能够更有效地开发它们在下游应用中的使用。在这项工作中,我们开始填补这一空白,通过使用一套分析工具,通过其中间表示向量,检查一个最近成功的预训练模型(Wav2vec 2.0)。我们使用典型相关性、互信息和具有非参数探测器的简单下游任务的性能的度量,以便(I)查询声学和语言信息内容,(Ii)表征跨模型层的信息的演变,以及(Iii)了解自动语音识别(ASR)的模型微调如何影响这些观测。我们的发现促使我们修改ASR的微调协议,在低资源环境下产生更高的单词错误率。
Recently proposed self-supervised learning approaches have been successful for pre-training speech representation models. The utility of these learned representations has been observed empirically, but not much has been studied about the type or extent of information encoded in the pre-trained representations themselves. Developing such insights can help understand the capabilities and limits of these models and enable the research community to more efficiently develop their usage for downstream applications. In this work, we begin to fill this gap by examining one recent and successful pre-trained model (wav2vec 2.0), via its intermediate representation vectors, using a suite of analysis tools. We use the metrics of canonical correlation, mutual information, and performance on simple downstream tasks with non-parametric probes, in order to (i) query for acoustic and linguistic information content, (ii) characterize the evolution of information across model layers, and (iii) understand how fine-tuning the model for automatic speech recognition (ASR) affects these observations. Our findings motivate modifying the fine-tuning protocol for ASR, which produces improved word error rates in a low-resource setting.