Feudal Reinforcement Learning for Dialogue Management in Large Domains

Feudal Reinforcement Learning for Dialogue Management in Large Domains
复制标题

DOI:
10.18653/v1/n18-2112
复制
发表时间:
2018-03
期刊:
ArXiv
影响因子:
--
通讯作者:
I. Casanueva;Paweł Budzianowski;Pei-hao Su;Stefan Ultes;L. Rojas-Barahona;Bo-Hsiang Tseng;M. Gašić
I. Casanueva;Paweł Budzianowski;Pei-hao Su;Stefan Ultes;L. Rojas-Barahona;Bo-Hsiang Tseng;M. Gašić
中科院分区:
其他
文献类型:
--
作者:
I. Casanueva;Paweł Budzianowski;Pei-hao Su;Stefan Ultes;L. Rojas-Barahona;Bo-Hsiang Tseng;M. Gašić

文献摘要

被引文献

相似文献

强化学习(RL)是解决对话策略优化的一种有前途的方法。然而,由于维数灾难,传统的强化学习算法无法扩展到大域。我们提出了一种基于 Feudal RL 的新型对话管理架构,它将决策分解为两个步骤;第一步,主策略选择原始动作的子集;第二步,从所选子集中选择原始动作。领域本体中包含的结构信息用于抽象对话状态空间,使用抽象状态的不同部分在每一步做出决策。这与时隙之间的信息共享机制相结合,提高了大域的可扩展性。我们证明,这种基于 Deep-Q 网络的方法的实现在多个对话领域和环境中显着优于以前的最先进技术,并且不需要任何额外的奖励信号。
Reinforcement learning (RL) is a promising approach to solve dialogue policy optimisation. Traditional RL algorithms, however, fail to scale to large domains due to the curse of dimensionality. We propose a novel Dialogue Management architecture, based on Feudal RL, which decomposes the decision into two steps; a first step where a master policy selects a subset of primitive actions, and a second step where a primitive action is chosen from the selected subset. The structural information included in the domain ontology is used to abstract the dialogue state space, taking the decisions at each step using different parts of the abstracted state. This, combined with an information sharing mechanism between slots, increases the scalability to large domains. We show that an implementation of this approach, based on Deep-Q Networks, significantly outperforms previous state of the art in several dialogue domains and environments, without the need of any additional reward signal.