Reinforcement learning for dialog management using least-squares Policy iteration and fast feature selection

Reinforcement learning for dialog management using least-squares Policy iteration and fast feature selection
复制标题

DOI:
10.21437/interspeech.2009-659
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
Lihong Li;J. Williams;Suhrid Balakrishnan
Lihong Li;J. Williams;Suhrid Balakrishnan
中科院分区:
其他
文献类型:
--
作者:
Lihong Li;J. Williams;Suhrid Balakrishnan

文献摘要

被引文献

相似文献

强化学习(RL)是一种很有前途的创建对话管理器的技术。RL接受当前对话框状态的功能,并寻求找到给定这些功能的最佳操作。虽然它往往是很容易的潜在有用的功能,在实践中,它是很难找到的子集是足够大,包含有用的信息,但足够紧凑,可靠地学习一个好的策略。在本文中,我们提出了一种自动执行特征选择的RL优化方法。该算法是基于最小二乘策略迭代,一个国家的最先进的强化学习算法,这是高样本效率,可以从静态语料库或在线学习。对话模拟实验表明,它比从一个工作的对话系统中提取的基线RL算法更稳定。
Reinforcement learning (RL) is a promising technique for creating a dialog manager. RL accepts features of the current dialog state and seeks to find the best action given those features. Although it is often easy to posit a large set of potentially useful features, in practice, it is difficult to find the subset which is large enough to contain useful information yet compact enough to reliably learn a good policy. In this paper, we propose a method for RL optimization which automatically performs feature selection. The algorithm is based on least-squares policy iteration, a state-of-the-art RL algorithm which is highly sampleefficient and can learn from a static corpus or on-line. Experiments in dialog simulation show it is more stable than a baseline RL algorithm taken from a working dialog system.