Training Spoken Language Understanding Systems with Non-Parallel Speech and Text

Training Spoken Language Understanding Systems with Non-Parallel Speech and Text
复制标题

DOI:
10.1109/icassp40776.2020.9054664
复制
发表时间:
2020-05
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Leda Sari;Samuel Thomas;M. Hasegawa-Johnson
Leda Sari;Samuel Thomas;M. Hasegawa-Johnson
中科院分区:
其他
文献类型:
--
作者:
Leda Sari;Samuel Thomas;M. Hasegawa-Johnson

文献摘要

相似文献

端到端口语理解(SLU)系统通常是在大量数据上训练的。在许多实际场景中,与文本相反,标记语音的量通常是有限的。在这项研究中,我们研究了使用非平行的语音和文本,以提高性能的对话行为识别作为一个例子SLU任务。我们提出了一个多视图架构,可以单独处理每种形式。为了有效地对这些数据进行训练,该模型使用共享分类器强制内部语音和文本编码相似。在Switchboard Dialog Act语料库上,我们表明使用大量文本预训练分类器有助于学习更好的语音编码,从而获得高达40%的相对较高的分类准确率。我们还表明,当语音嵌入自动语音识别(ASR)系统中使用在这个框架中,语音的准确性超过了ASR的基于文本的测试高达15%的相对性能,并接近使用真实成绩单的性能。
End-to-end spoken language understanding (SLU) systems are typically trained on large amounts of data. In many practical scenarios, the amount of labeled speech is often limited as opposed to text. In this study, we investigate the use of non-parallel speech and text to improve the performance of dialog act recognition as an example SLU task. We propose a multiview architecture that can handle each modality separately. To effectively train on such data, this model enforces the internal speech and text encodings to be similar using a shared classifier. On the Switchboard Dialog Act corpus, we show that pretraining the classifier using large amounts of text helps learning better speech encodings, resulting in up to 40% relatively higher classification accuracies. We also show that when the speech embeddings from an automatic speech recognition (ASR) system are used in this framework, the speech-only accuracy exceeds the performance of ASR-text based tests up to 15% relative and approaches the performance of using true transcripts.