RNN Transducer Models for Spoken Language Understanding

RNN Transducer Models for Spoken Language Understanding
复制标题

用于口语理解的 RNN 换能器模型

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
R. Hoory
R. Hoory
中科院分区:
--
文献类型:
--
作者:
Samuel Thomas;H. Kuo;G. Saon;Zolt'an Tuske;Brian Kingsbury;Gakuto Kurata;Zvi Kons;R. Hoory

文献摘要

被引文献

相似文献

我们提出了一个全面的研究建立和适应RNN转换器(RNN-T)模型的口语理解(SLU)。这些端到端(E2 E)模型构建在三个实际的设置:一个逐字记录的情况下,可用,一个约束的情况下,唯一可用的注释是SLU标签和它们的值,和一个更严格的情况下,转录可用,但没有相应的音频。我们展示了如何从预训练的自动语音识别(ASR)系统开始开发RNN-T SLU模型,然后进行SLU自适应步骤。在真实的音频数据不可用的设置中,使用人工合成的语音来成功地适应各种SLU模型。当评估两个SLU数据集,ATIS语料库和客户呼叫中心数据集,所提出的模型密切跟踪其他E2 E模型的性能,并实现国家的最先进的结果。
We present a comprehensive study on building and adapting RNN transducer (RNN-T) models for spoken language understanding (SLU). These end-to-end (E2E) models are constructed in three practical settings: a case where verbatim transcripts are available, a constrained case where the only available annotations are SLU labels and their values, and a more restrictive case where transcripts are available but not corresponding audio. We show how RNN-T SLU models can be developed starting from pre-trained automatic speech recognition (ASR) systems, followed by an SLU adaptation step. In settings where real audio data is not available, artificially synthesized speech is used to successfully adapt various SLU models. When evaluated on two SLU data sets, the ATIS corpus and a customer call center data set, the proposed models closely track the performance of other E2E models and achieve state-of-the-art results.