Model-based reinforcement learning with dimension reduction
Model-based reinforcement learning with dimension reduction
复制标题
DOI:
10.1016/j.neunet.2016.08.005
复制
发表时间:
2016-12
期刊:
影响因子:
--
通讯作者:
Voot Tangkaratt;J. Morimoto;Masashi Sugiyama
中科院分区:
文献类型:
--
作者:
Voot Tangkaratt;J. Morimoto;Masashi Sugiyama
The goal of reinforcement learning is to learn an optimal policy which controls an agent to acquire the maximum cumulative reward. Themodel-basedreinforcement learning approach learns a transition model of the environment from data, and then derives the optimal policy using the transition model. However, learning an accurate transition model in high-dimensional environments requires a large amount of data which is difficult to obtain. To overcome this difficulty, in this paper, we propose to combine model-based reinforcement learning with the recently developedleast-squares conditional entropy(LSCE) method, which simultaneously performs transition model estimation and dimension reduction. We also further extend the proposed method to imitation learning scenarios. The experimental results show that policy search combined with LSCE performs well for high-dimensional control tasks including real humanoid robot control.