Reward-Weighted Regression with Sample Reuse for Direct Policy Search in Reinforcement Learning

Reward-Weighted Regression with Sample Reuse for Direct Policy Search in Reinforcement Learning
复制标题

DOI:
10.1162/neco_a_00199
复制
发表时间:
2011-11
期刊:
影响因子:
2.9
通讯作者:
Hirotaka Hachiya;Jan Peters;Masashi Sugiyama
Hirotaka Hachiya;Jan Peters;Masashi Sugiyama
中科院分区:
计算机科学4区
文献类型:
--
作者:
Hirotaka Hachiya;Jan Peters;Masashi Sugiyama

文献摘要

相似文献

摘要直接策略搜索是一种很有前途的强化学习框架,特别适用于控制连续的高维系统。策略搜索通常需要大量的样本来获得稳定的策略更新估计器,当采样成本昂贵时,这是令人望而却步的。在这封信中,我们扩展了一种基于期望最大化的策略搜索方法,使得以前收集的样本可以有效地重复使用。通过机器人学习实验,验证了基于样本重复使用的奖励加权回归方法的有效性。(这封信是我们之前会议论文的扩展版本:Hchiya,Peters,&Sugiyama,2009。)
Abstract Direct policy search is a promising reinforcement learning framework, in particular for controlling continuous, high-dimensional systems. Policy search often requires a large number of samples for obtaining a stable policy update estimator, and this is prohibitive when the sampling cost is expensive. In this letter, we extend an expectation-maximization-based policy search method so that previously collected samples can be efficiently reused. The usefulness of the proposed method, reward-weighted regression with sample reuse (R), is demonstrated through robot learning experiments. (This letter is an extended version of our earlier conference paper: Hachiya, Peters, & Sugiyama, 2009.)