Adaptive Dialog Policy Learning with Hindsight and User Modeling

Adaptive Dialog Policy Learning with Hindsight and User Modeling
复制标题

DOI:
10.18653/v1/2020.sigdial-1.40
复制
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yan Cao;Keting Lu;Xiaoping Chen;Shiqi Zhang
Yan Cao;Keting Lu;Xiaoping Chen;Shiqi Zhang
中科院分区:
其他
文献类型:
--
作者:
Yan Cao;Keting Lu;Xiaoping Chen;Shiqi Zhang

文献摘要

被引文献

相似文献

强化学习(RL)方法已被广泛用于学习对话策略。样本效率,即从有限的对话经验中学习的效率,在基于RL的对话策略学习中尤为重要,因为与人的互动是昂贵且低质量的对话策略会产生非常差的用户体验。在本文中,我们开发了LHUA(以事后看来,用户建模和适应性学习),这首先使对话框代理可以从模拟和真实用户中的事后自适应地学习。模拟和事后看来,对话框分别提供了更多的经验和更多(积极的)强化。实验结果表明,LHUA的表现优于文献中的竞争基准,包括其无仿真,无适应和无视力。
Reinforcement learning (RL) methods have been widely used for learning dialog policies. Sample efficiency, i.e., the efficiency of learning from limited dialog experience, is particularly important in RL-based dialog policy learning, because interacting with people is costly and low-quality dialog policies produce very poor user experience. In this paper, we develop LHUA (Learning with Hindsight, User modeling, and Adaptation) that, for the first time, enables dialog agents to adaptively learn with hindsight from both simulated and real users. Simulation and hindsight provide the dialog agent with more experience and more (positive) reinforcement respectively. Experimental results suggest that LHUA outperforms competitive baselines from the literature, including its no-simulation, no-adaptation, and no-hindsight counterparts.