Hybrid Reinforcement/Supervised Learning of Dialogue Policies from Fixed Data Sets

Hybrid Reinforcement/Supervised Learning of Dialogue Policies from Fixed Data Sets
复制标题

DOI:
10.1162/coli.2008.07-028-r2-05-82
复制
发表时间:
2008-12
影响因子:
9.3
通讯作者:
James Henderson;Oliver Lemon;Kallirroi Georgila
James Henderson;Oliver Lemon;Kallirroi Georgila
中科院分区:
计算机科学3区
文献类型:
--
作者:
James Henderson;Oliver Lemon;Kallirroi Georgila

文献摘要

被引文献

相似文献

摘要我们提出了一种从固定数据集学习对话管理策略的方法。该方法解决了基于信息状态更新(ISU)的对话系统,它代表了一个对话的状态作为一个大的功能集,导致一个非常大的状态空间和一个巨大的政策空间所带来的挑战。为了解决任何固定数据集只提供这些状态和策略空间的一小部分信息的问题,我们提出了一种将强化学习与监督学习相结合的混合模型。强化学习用于优化对话奖励的度量,而监督学习用于将学习的策略限制在我们有数据的这些空间的部分。我们还使用线性函数近似来解决从固定数量的数据推广到大状态空间的需求。为了证明这种方法在这个具有挑战性的任务上的有效性,我们在COMMUNICATOR语料库上训练了这个模型,我们为用户操作和信息状态添加了注释。当使用在同一数据集的不同部分上训练的用户模拟进行测试时,我们的混合模型优于纯监督学习模型和纯强化学习模型。根据自动评估措施,它在COMMUNICATOR数据上的性能还优于手工制作的系统,比平均COMMUNICATOR系统策略提高了10%。所提出的方法将改善技术的引导和自动优化的对话管理政策,从有限的初始数据集。
Abstract We propose a method for learning dialogue management policies from a fixed data set. The method addresses the challenges posed by Information State Update (ISU)-based dialogue systems, which represent the state of a dialogue as a large set of features, resulting in a very large state space and a huge policy space. To address the problem that any fixed data set will only provide information about small portions of these state and policy spaces, we propose a hybrid model that combines reinforcement learning with supervised learning. The reinforcement learning is used to optimize a measure of dialogue reward, while the supervised learning is used to restrict the learned policy to the portions of these spaces for which we have data. We also use linear function approximation to address the need to generalize from a fixed amount of data to large state spaces. To demonstrate the effectiveness of this method on this challenging task, we trained this model on the COMMUNICATOR corpus, to which we have added annotations for user actions and Information States. When tested with a user simulation trained on a different part of the same data set, our hybrid model outperforms a pure supervised learning model and a pure reinforcement learning model. It also outperforms the hand-crafted systems on the COMMUNICATOR data, according to automatic evaluation measures, improving over the average COMMUNICATOR system policy by 10%. The proposed method will improve techniques for bootstrapping and automatic optimization of dialogue management policies from limited initial data sets.