Using trajectory data to improve bayesian optimization for reinforcement learning

Using trajectory data to improve bayesian optimization for reinforcement learning
复制标题

DOI:
10.5555/2627435.2627443
复制
发表时间:
2014
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Aaron Wilson;Alan Fern;Prasad Tadepalli
Aaron Wilson;Alan Fern;Prasad Tadepalli
中科院分区:
其他
文献类型:
--
作者:
Aaron Wilson;Alan Fern;Prasad Tadepalli

文献摘要

被引文献

相似文献

最近,贝叶斯优化(BO)已被成功地用于在几个具有挑战性的强化学习(RL)应用中优化参数策略。BO对这个问题很有吸引力,因为它利用了关于预期收益的贝叶斯先验信息,并利用这些信息来选择要执行的新策略。有效地,用于政策搜索的BO框架解决了勘探和开采之间的权衡。在这项工作中,我们展示了如何通过利用RL代理生成的序列轨迹信息更有效地将BO应用于RL。我们的贡献可以分为两个截然不同但互惠互利的部分。第一种是一种新的高斯过程(GP)核,用于使用策略执行产生的轨迹数据来衡量策略之间的相似性。该核函数可用于改进预期收益的后验估计,从而提高勘探质量。第二个贡献是一种新的GP均值函数,它使用学习的转移函数和奖励函数来逼近目标的表面。我们表明,我们开发的基于模型的方法可以从模型的不准确中恢复,当无法学习良好的过渡和回报模型时。我们在一组标准的RL基准测试中给出的实验结果表明,与竞争方法相比,我们的基于模型和无模型的方法都可以加快学习速度。此外,我们表明,我们的贡献可以结合在一起,在某些领域产生协同改进。
Recently, Bayesian Optimization (BO) has been used to successfully optimize parametric policies in several challenging Reinforcement Learning (RL) applications. BO is attractive for this problem because it exploits Bayesian prior information about the expected return and exploits this knowledge to select new policies to execute. Effectively, the BO framework for policy search addresses the exploration-exploitation tradeoff. In this work, we show how to more effectively apply BO to RL by exploiting the sequential trajectory information generated by RL agents. Our contributions can be broken into two distinct, but mutually beneficial, parts. The first is a new Gaussian process (GP) kernel for measuring the similarity between policies using trajectory data generated from policy executions. This kernel can be used in order to improve posterior estimates of the expected return thereby improving the quality of exploration. The second contribution, is a new GP mean function which uses learned transition and reward functions to approximate the surface of the objective. We show that the model-based approach we develop can recover from model inaccuracies when good transition and reward models cannot be learned. We give empirical results in a standard set of RL benchmarks showing that both our model-based and model-free approaches can speed up learning compared to competing methods. Further, we show that our contributions can be combined to yield synergistic improvement in some domains.