Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent Systems

Tokenwise Contrastive Pretraining for Finer Speech-to-BERT Alignment in End-to-End Speech-to-Intent Systems
复制标题

用于端到端语音到意图系统中更精细的语音到 BERT 对齐的 Tokenwise 对比预训练

DOI:
--
复制
发表时间:
2022
期刊:
Interspeech
影响因子:
--
通讯作者:
Brian Kingsbury
Brian Kingsbury
中科院分区:
--
文献类型:
--
作者:
Vishal Sunder;E. Fosler;Samuel Thomas;H. Kuo;Brian Kingsbury

文献摘要

被引文献

相似文献

端到端 (E2E) 口语理解 (SLU) 的最新进展主要归功于语音表示的有效预训练。其中一种预训练范式是将语义知识从最先进的基于文本的模型(如 BERT)提炼到语音编码器神经网络。这项工作是朝着以更高效、更细粒度的方式做同样的事情迈出的一步,我们在逐个标记的基础上对齐语音嵌入和 BERT 嵌入。我们引入了一种简单而新颖的技术,该技术使用跨模式注意机制从语音编码器中提取令牌级上下文嵌入,以便可以将它们直接与基于 BERT 的上下文嵌入进行比较和对齐。这种对齐是使用一种新颖的标记对比损失来执行的。对这样的预训练模型进行微调,以直接使用语音执行意图识别,从而在两个广泛使用的 SLU 数据集上产生最先进的性能。当使用 SpecAugment 进行额外的正则化微调时,我们的模型得到了进一步的改进,特别是当语音有噪声时,与之前的结果相比,绝对改进高达 8%。
Recent advances in End-to-End (E2E) Spoken Language Understanding (SLU) have been primarily due to effective pretraining of speech representations. One such pretraining paradigm is the distillation of semantic knowledge from state-of-the-art text-based models like BERT to speech encoder neural networks. This work is a step towards doing the same in a much more efficient and fine-grained manner where we align speech embeddings and BERT embeddings on a token-by-token basis. We introduce a simple yet novel technique that uses a cross-modal attention mechanism to extract token-level contextual embeddings from a speech encoder such that these can be directly compared and aligned with BERT based contextual embeddings. This alignment is performed using a novel tokenwise contrastive loss. Fine-tuning such a pretrained model to perform intent recognition using speech directly yields state-of-the-art performance on two widely used SLU datasets. Our model improves further when fine-tuned with additional regularization using SpecAugment especially when speech is noisy, giving an absolute improvement as high as 8% over previous results.