Model-based reinforcement learning with dimension reduction

Model-based reinforcement learning with dimension reduction
复制标题

DOI:
10.1016/j.neunet.2016.08.005
复制
发表时间:
2016-12
期刊:
Neural networks : the official journal of the International Neural Network Society
影响因子:
--
通讯作者:
Voot Tangkaratt;J. Morimoto;Masashi Sugiyama
Voot Tangkaratt;J. Morimoto;Masashi Sugiyama
中科院分区:
其他
文献类型:
--
作者:
Voot Tangkaratt;J. Morimoto;Masashi Sugiyama

文献摘要

被引文献

相似文献

强化学习的目标是学习一个最优策略,控制智能体获得最大的累积奖励。基于模型的强化学习方法从数据中学习环境的过渡模型,然后使用过渡模型导出最优策略。然而,在高维环境中学习准确的过渡模型需要大量的数据,这是很难获得的。为了克服这个困难,在本文中,我们提出了联合收割机基于模型的强化学习与最近开发的最小二乘条件熵(LSCE)方法,同时进行过渡模型估计和降维。我们还进一步扩展所提出的方法模仿学习的情况下。实验结果表明,结合LSCE的策略搜索对于包括真实的仿人机器人控制在内的高维控制任务表现良好。
The goal of reinforcement learning is to learn an optimal policy which controls an agent to acquire the maximum cumulative reward. Themodel-basedreinforcement learning approach learns a transition model of the environment from data, and then derives the optimal policy using the transition model. However, learning an accurate transition model in high-dimensional environments requires a large amount of data which is difficult to obtain. To overcome this difficulty, in this paper, we propose to combine model-based reinforcement learning with the recently developedleast-squares conditional entropy(LSCE) method, which simultaneously performs transition model estimation and dimension reduction. We also further extend the proposed method to imitation learning scenarios. The experimental results show that policy search combined with LSCE performs well for high-dimensional control tasks including real humanoid robot control.