End-to-End Spoken Language Understanding Without Full Transcripts

End-to-End Spoken Language Understanding Without Full Transcripts
复制标题

DOI:
10.21437/interspeech.2020-2924
复制
发表时间:
2020-09
期刊:
ArXiv
影响因子:
--
通讯作者:
H. Kuo;Zolt'an Tuske;Samuel Thomas;Yinghui Huang;Kartik Audhkhasi;Brian Kingsbury;Gakuto Kurata;Zvi Kons;R. Hoory;L. Lastras
H. Kuo;Zolt'an Tuske;Samuel Thomas;Yinghui Huang;Kartik Audhkhasi;Brian Kingsbury;Gakuto Kurata;Zvi Kons;R. Hoory;L. Lastras
中科院分区:
其他
文献类型:
--
作者:
H. Kuo;Zolt'an Tuske;Samuel Thomas;Yinghui Huang;Kartik Audhkhasi;Brian Kingsbury;Gakuto Kurata;Zvi Kons;R. Hoory;L. Lastras

文献摘要

被引文献

相似文献

口语理解 (SLU) 的一个重要组成部分是槽填充:使用语义实体标签表示口语表达的含义。在本文中,我们开发了端到端 (E2E) 口语理解系统,可直接将语音输入转换为语义实体,并研究这些 E2E SLU 模型是否可以仅在语义实体注释上进行训练,而无需逐字转录。训练此类模型非常有用,因为它们可以大大降低数据收集的成本。我们通过调整最初为语音识别训练的模型,创建了两种类型的语音到实体模型:CTC 模型和基于注意力的编码器-解码器模型。鉴于我们的实验涉及语音输入,这些系统需要正确识别实体标签和代表实体值的单词。对于我们在 ATIS 语料库上进行的语音到实体的实验,CTC 和注意力模型都显示出令人印象深刻的跳过非实体单词的能力:仅对实体进行训练与对完整转录本进行训练时几乎没有退化。我们还探讨了实体的顺序不一定与话语中的语音顺序相关的场景。凭借其重新排序的能力,注意力模型表现非常出色,语音到实体袋的 F1 分数仅下降了约 2%。
An essential component of spoken language understanding (SLU) is slot filling: representing the meaning of a spoken utterance using semantic entity labels. In this paper, we develop end-to-end (E2E) spoken language understanding systems that directly convert speech input to semantic entities and investigate if these E2E SLU models can be trained solely on semantic entity annotations without word-for-word transcripts. Training such models is very useful as they can drastically reduce the cost of data collection. We created two types of such speech-to-entities models, a CTC model and an attention-based encoder-decoder model, by adapting models trained originally for speech recognition. Given that our experiments involve speech input, these systems need to recognize both the entity label and words representing the entity value correctly. For our speech-to-entities experiments on the ATIS corpus, both the CTC and attention models showed impressive ability to skip non-entity words: there was little degradation when trained on just entities versus full transcripts. We also explored the scenario where the entities are in an order not necessarily related to spoken order in the utterance. With its ability to do re-ordering, the attention model did remarkably well, achieving only about 2% degradation in speech-to-bag-of-entities F1 score.