Towards integrated dialogue policy learning for multiple domains and intents using Hierarchical Deep Reinforcement Learning

Towards integrated dialogue policy learning for multiple domains and intents using Hierarchical Deep Reinforcement Learning
复制标题

DOI:
10.1016/j.eswa.2020.113650
复制
发表时间:
2020-12-30
影响因子:
8.5
通讯作者:
Bhattacharyya, Pushpak
Bhattacharyya, Pushpak
中科院分区:
计算机科学1区
文献类型:
--
作者:
Saha, Tulika;Gupta, Dhawal;Bhattacharyya, Pushpak

文献摘要

被引文献

相似文献

创建专家和智能对话/虚拟代理(VA),可以服务于与多个域及其各种意图相关的复杂和复杂的用户任务(需求),确实是相当具有挑战性的,因为它需要代理同时处理不同域中的多个子任务。本文提出了一个专家的、统一的和通用的深度强化学习(DRL)框架,它创建了对话管理器,能够管理包含多个领域及其各种意图的面向任务的对话,并为用户提供一个一站式的专家系统。为了解决这多个方面,用户和VA之间的对话交换被分成多个层次,以便隔离和识别属于不同域的子任务。分层强化学习(HRL)的概念特别是选项被用来学习这些分层结构中的最优策略,这些策略在不同的时间步长上操作以实现用户目标。对话管理器包括顶层域元策略、中级意图元策略,以便在各种和多个子任务或选项中进行选择,以及低级控制器策略,以选择基本动作以完成由不同意图和域中的较高层元策略给出的子任务。在重叠子任务之间共享控制器策略使得元策略成为通用的。拟议的专家框架已在“航空旅行”和“餐馆”领域进行了演示。与几个强有力的基线和最先进的模型相比,实验确定了学习的政策的效率,以及对能够处理复杂和复合任务的这种专家模型的需求。(C)2020爱思唯尔有限公司。保留所有权利。
Creation of Expert and Intelligent Dialogue/Virtual Agent (VA) that can serve complicated and intricate tasks (need) of the user related to multiple domains and its various intents is indeed quite challenging as it necessitates the agent to concurrently handle multiple subtasks in different domains. This paper presents an expert, unified and a generic Deep Reinforcement Learning (DRL) framework that creates dialogue managers competent for managing task-oriented conversations embodying multiple domains along with their various intents and provide the user with an expert system which is a one stop for all queries. In order to address these multiple aspects, the dialogue exchange between the user and the VA is split into hierarchies, so as to isolate and identify subtasks belonging to different domains. The notion of Hierarchical Reinforcement Learning (HRL) specifically options is employed to learn optimal policies in these hierarchies that operate at varying time steps to accomplish the user goal. The dialogue manager encompasses a top-level domain meta-policy, intermediate-level intent meta-policies in order to select amongst varied and multiple subtasks or options and low-level controller policies to select primitive actions to complete the subtask given by the higher-level meta-policies in varying intents and domains. Sharing of controller policies among overlapping subtasks enables the meta-policies to be generic. The proposed expert framework has been demonstrated in the domains of "Air Travel" and "Restaurant". Experiments as compared to several strong baselines and a state of the art model establish the efficiency of the learned policies and the need for such expert models capable of handling complex and composite tasks. (C) 2020 Elsevier Ltd. All rights reserved.