Lattice-Based Improvements for Voice Triggering Using Graph Neural Networks

Lattice-Based Improvements for Voice Triggering Using Graph Neural Networks
复制标题

使用图神经网络对语音触发进行基于格的改进

DOI:
--
复制
发表时间:
2020
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Jason D. Williams
Jason D. Williams
中科院分区:
--
文献类型:
--
作者:
Pranay Dighe;Saurabh N. Adya;Nuoyu Li;Srikanth Vishnubhotla;D. Naik;Adithya Sagar;Ying Ma;S. Pulman;Jason D. Williams

文献摘要

参考文献

被引文献

相似文献

语音触发的智能助理通常依赖于在开始监听用户请求之前检测到一个短语。减少错误触发是构建以隐私为中心的非侵入式智能助理的一个重要方面。在本文中,我们解决了错误触发缓解(FTM)的任务,使用一种新的方法的基础上分析自动语音识别(ASR)格使用图神经网络(GNN)。所提出的方法使用的事实是,解码晶格的错误触发的音频表现出不确定性方面的许多替代路径和意想不到的词的晶格弧相比,晶格的正确触发的音频。纯粹的短语检测器模型不能充分利用用户语音的意图,而通过使用用户音频的完整解码网格,我们可以有效地减轻不针对智能助理的语音。在本文中,我们分别基于1)图卷积层和2)自注意机制部署了两种GNN变体。我们的实验表明,GNN在FTM任务中具有高度的准确性,在99%的真阳性率(TPR)下减少了约87%的错误触发。此外,所提出的模型是快速训练和有效的参数要求。
Voice-triggered smart assistants often rely on detection of a trigger-phrase before they start listening for the user request. Mitigation of false triggers is an important aspect of building a privacy-centric non-intrusive smart assistant. In this paper, we address the task of false trigger mitigation (FTM) using a novel approach based on analyzing automatic speech recognition (ASR) lattices using graph neural networks (GNN). The proposed approach uses the fact that decoding lattice of a falsely triggered audio exhibits uncertainties in terms of many alternative paths and unexpected words on the lattice arcs as compared to the lattice of a correctly triggered audio. A pure trigger-phrase detector model doesn’t fully utilize the intent of the user speech whereas by using the complete decoding lattice of user audio, we can effectively mitigate speech not intended for the smart assistant. We deploy two variants of GNNs in this paper based on 1) graph convolution layers and 2) self-attention mechanism respectively. Our experiments demonstrate that GNNs are highly accurate in FTM task by mitigating ~87% of false triggers at 99% true positive rate (TPR). Furthermore, the proposed models are fast to train and efficient in parameter requirements.
DOI: 10.1109/tnnls.2020.2978386
发表时间: 2021-01-01
影响因子: 10.4
作者:
Wu, Zonghan;Pan, Shirui;Yu, Philip S.
通讯作者: Yu, Philip S.