Integration of Pre-Trained Networks with Continuous Token Interface for End-to-End Spoken Language Understanding

Integration of Pre-Trained Networks with Continuous Token Interface for End-to-End Spoken Language Understanding
复制标题

DOI:
10.1109/icassp43922.2022.9747047
复制
发表时间:
2021-04
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
S. Seo;Donghyun Kwak;Bowon Lee
S. Seo;Donghyun Kwak;Bowon Lee
中科院分区:
其他
文献类型:
--
作者:
S. Seo;Donghyun Kwak;Bowon Lee

文献摘要

被引文献

相似文献

大多数端到端 (E2E) 口语理解 (SLU) 网络利用预先训练的自动语音识别 (ASR) 网络,但仍然缺乏理解话语语义的能力,而这对于 SLU 任务至关重要。为了解决这个问题,最近提出的研究使用预先训练的自然语言理解(NLU)网络。然而,充分利用这两个预训练网络并非易事。提出了许多解决方案,例如知识蒸馏(KD)、跨模式共享嵌入以及与接口的网络集成。我们提出了一种简单而鲁棒的 E2E SLU 网络集成方法,具有新颖的接口,连续令牌接口(CTI)。当 ASR 和 NLU 网络都使用相同的词汇进行预训练时,CTI 是 ASR 和 NLU 网络的连接表示。因此,我们可以以端到端的方式训练我们的 SLU 网络,而无需额外的模块,例如 Gumbel-Softmax。我们使用 SLURP(一个具有挑战性的 SLU 数据集)评估我们的模型,并在意图分类和槽填充任务上获得最先进的分数。我们还验证了使用掩码语言模型 (MLM) 进行预训练的 NLU 网络可以利用 CTI 的噪声文本表示。此外,我们使用额外的数据 SLURP-Synth 来训练我们的模型,并获得了更好的结果。
Most End-to-End (E2E) Spoken Language Understanding (SLU) networks leverage the pre-trained Automatic Speech Recognition (ASR) networks but still lack the capability to understand the semantics of utterances, crucial for the SLU task. To solve this, recently proposed studies use pre-trained Natural Language Understanding (NLU) networks. However, it is not trivial to fully utilize both pre-trained networks; many solutions were proposed, such as Knowledge Distillation (KD), cross-modal shared embedding and network integration with Interface. We propose a simple and robust integration method for the E2E SLU network with a novel Interface, Continuous Token Interface (CTI). CTI is a junctional representation of the ASR and NLU networks when both networks are pre-trained with the same vocabulary. Thus, we can train our SLU network in an E2E manner without additional modules, such as Gumbel-Softmax. We evaluate our model using SLURP, a challenging SLU dataset and achieve state-of-the-art scores on intent classification and slot filling tasks. We also verify that the NLU network, pre-trained with Masked Language Model (MLM), can utilize a noisy textual representation of CTI. Moreover, we train our model with extra data, SLURP-Synth, and get better results.