InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning

InstructDial: Improving Zero and Few-shot Generalization in Dialogue through Instruction Tuning
复制标题

DOI:
10.18653/v1/2022.emnlp-main.33
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Prakhar Gupta;Cathy Jiao;Yi-Ting Yeh;Shikib Mehri;M. Eskénazi;Jeffrey P. Bigham
Prakhar Gupta;Cathy Jiao;Yi-Ting Yeh;Shikib Mehri;M. Eskénazi;Jeffrey P. Bigham
中科院分区:
其他
文献类型:
--
作者:
Prakhar Gupta;Cathy Jiao;Yi-Ting Yeh;Shikib Mehri;M. Eskénazi;Jeffrey P. Bigham

文献摘要

被引文献

相似文献

指令微调是自然语言处理(NLP)中一种新兴的范式,其中自然语言指令与语言模型相结合,以在未见过的任务上诱导零样本性能。对话是探索指令微调特别有趣的一个领域,因为对话系统执行多种与语言相关的任务(例如,自然语言理解和生成、特定领域的交互),然而指令微调尚未针对对话相关任务进行系统地探索。我们引入了InstructDial,一个用于对话的指令微调框架,它由一个包含48个不同对话任务的知识库组成,这些任务以统一的文本到文本格式从59个公开可用的对话数据集中创建。我们探索了在InstructDial上微调的模型在不同对话任务上的跨任务泛化能力。我们的分析表明,InstructDial能够在未见过的数据集和任务(如对话评估和意图检测)上实现良好的零样本性能,在少样本设置下甚至有更好的性能。为了确保模型遵循指令,我们引入了新的元任务。我们确定了使用所提出的框架在多个对话任务上训练的模型的基准零样本和少样本性能。
Instruction tuning is an emergent paradigm in NLP wherein natural language instructions are leveraged with language models to induce zero-shot performance on unseen tasks. Dialogue is an especially interesting area in which to explore instruction tuning because dialogue systems perform multiple kinds of tasks related to language (e.g., natural language understanding and generation, domain-specific interaction), yet instruction tuning has not been systematically explored for dialogue-related tasks. We introduce InstructDial, an instruction tuning framework for dialogue, which consists of a repository of 48 diverse dialogue tasks in a unified text-to-text format created from 59 openly available dialogue datasets. We explore cross-task generalization ability on models tuned on InstructDial across diverse dialogue tasks. Our analysis reveals that InstructDial enables good zero-shot performance on unseen datasets and tasks such as dialogue evaluation and intent detection, and even better performance in a few-shot setting. To ensure that models adhere to instructions, we introduce novel meta-tasks. We establish benchmark zero-shot and few-shot performance of models trained using the proposed framework on multiple dialogue tasks.