Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference

Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference
复制标题

DOI:
10.18653/v1/2020.emnlp-main.411
复制
发表时间:
2020-10
期刊:
--
影响因子:
--
通讯作者:
Jianguo Zhang;Kazuma Hashimoto;Wenhao Liu;Chien-Sheng Wu;Yao Wan;Philip S. Yu;R. Socher;Caiming Xiong
Jianguo Zhang;Kazuma Hashimoto;Wenhao Liu;Chien-Sheng Wu;Yao Wan;Philip S. Yu;R. Socher;Caiming Xiong
中科院分区:
其他
文献类型:
--
作者:
Jianguo Zhang;Kazuma Hashimoto;Wenhao Liu;Chien-Sheng Wu;Yao Wan;Philip S. Yu;R. Socher;Caiming Xiong

文献摘要

被引文献

相似文献

意图检测是目标导向对话系统的核心组成部分之一,而超范围意图检测也是一项重要的实践技能。Few-shot学习在缓解数据稀缺性方面吸引了很多关注,但OOS检测变得更具挑战性。本文提出了一种简单而有效的方法——基于深度自关注的判别最近邻分类。与softmax分类器不同,我们利用bert风格的成对编码来训练一个二元分类器,该分类器估计用户输入的最佳匹配训练示例。我们提出通过迁移自然语言推理(NLI)模型来提高识别能力。我们在大规模多领域意图检测任务上的大量实验表明,我们的方法比基于roberta的分类器和基于嵌入的最近邻方法实现了更稳定和准确的域内和OOS检测精度。更值得注意的是,NLI迁移使我们的10射击模型能够与50射击甚至全射击分类器竞争,同时我们可以通过利用更快的嵌入检索模型保持推理时间不变。
Intent detection is one of the core components of goal-oriented dialog systems, and detecting out-of-scope (OOS) intents is also a practically important skill. Few-shot learning is attracting much attention to mitigate data scarcity, but OOS detection becomes even more challenging. In this paper, we present a simple yet effective approach, discriminative nearest neighbor classification with deep self-attention. Unlike softmax classifiers, we leverage BERT-style pairwise encoding to train a binary classifier that estimates the best matched training example for a user input. We propose to boost the discriminative ability by transferring a natural language inference (NLI) model. Our extensive experiments on a large-scale multi-domain intent detection task show that our method achieves more stable and accurate in-domain and OOS detection accuracy than RoBERTa-based classifiers and embedding-based nearest neighbor approaches. More notably, the NLI transfer enables our 10-shot model to perform competitively with 50-shot or even full-shot classifiers, while we can keep the inference time constant by leveraging a faster embedding retrieval model.