What Do Audio Transformers Hear? Probing Their Representations For Language Delivery & Structure

What Do Audio Transformers Hear? Probing Their Representations For Language Delivery & Structure
复制标题

DOI:
10.1109/icdmw58026.2022.00120
复制
发表时间:
2022-11
期刊:
2022 IEEE International Conference on Data Mining Workshops (ICDMW)
影响因子:
--
通讯作者:
Yaman Kumar Singla;Jui Shah;Changyou Chen;R. Shah
Yaman Kumar Singla;Jui Shah;Changyou Chen;R. Shah
中科院分区:
其他
文献类型:
--
作者:
Yaman Kumar Singla;Jui Shah;Changyou Chen;R. Shah

文献摘要

相似文献

跨自然语言处理和语音等多个领域的Transformer模型是从业者和研究人员技术堆栈中不可避免的一部分。音频转换器利用表征学习来训练未标记的语音,最近已被用于从说话人确认到话语连贯性的任务,并取得了很大的成功。然而,人们对这些模型在高维潜在空间中学习和表示的内容知之甚少。在本文中,我们解释了两个这样的最近国家的最先进的模型,wav2vec2.0和Mockingjay,语言和声学特征。我们探索它们的每一层,以了解它在学习什么,同时,我们在两个模型之间进行了区分。通过比较他们在各种各样的设置,包括本机,非本机,阅读和自发演讲的性能,我们还显示了多少这些模型能够学习可转移的功能。我们的研究结果表明,该模型能够显着捕捉广泛的特征,如音频,流畅性,超音段发音,甚至句法和语义的文本为基础的特征。对于每一类特征,我们为每个框架确定一个学习模式,并得出结论,哪个模型和该模型的哪一层更适合于特定类别的特征,以选择用于下游任务的特征提取。
Transformer models across multiple domains such as natural language processing and speech form an unavoidable part of the tech stack of practitioners and researchers alike. Au-dio transformers that exploit representational learning to train on unlabeled speech have recently been used for tasks from speaker verification to discourse-coherence with much success. However, little is known about what these models learn and represent in the high-dimensional latent space. In this paper, we interpret two such recent state-of-the-art models, wav2vec2.0 and Mockingjay, on linguistic and acoustic features. We probe each of their layers to understand what it is learning and at the same time, we draw a distinction between the two models. By comparing their performance across a wide variety of settings including native, non-native, read and spontaneous speeches, we also show how much these models are able to learn transferable features. Our results show that the models are capable of significantly capturing a wide range of characteristics such as audio, fluency, supraseg-mental pronunciation, and even syntactic and semantic text-based characteristics. For each category of characteristics, we identify a learning pattern for each framework and conclude which model and which layer of that model is better for a specific category of feature to choose for feature extraction for downstream tasks.