Reinforcement Learning Based on On-Line EM Algorithm

Reinforcement Learning Based on On-Line EM Algorithm
复制标题

基于在线EM算法的强化学习

DOI:
--
复制
发表时间:
1998
期刊:
--
影响因子:
--
通讯作者:
S. Ishii
S. Ishii
中科院分区:
--
文献类型:
--
作者:
Masa;S. Ishii

文献摘要

被引文献

相似文献

在这篇文章中,我们提出了一种新的强化学习(RL)方法的基础上的演员-评论家架构。演员和评论家由归一化高斯网络(NGnet)近似,NGnet是局部线性回归单元的网络。NGnet的训练在线EM算法提出了我们以前的文件。我们应用我们的RL方法的摆动和稳定的单摆的任务和平衡的任务附近的直立位置的双摆。实验结果表明,我们的RL方法可以适用于具有连续状态/动作空间的最优控制问题,该方法实现了良好的控制与少量的试错。
In this article, we propose a new reinforcement learning (RL) method based on an actor-critic architecture. The actor and the critic are approximated by Normalized Gaussian Networks (NGnet), which are networks of local linear regression units. The NGnet is trained by the on-line EM algorithm proposed in our previous paper. We apply our RL method to the task of swinging-up and stabilizing a single pendulum and the task of balancing a double pendulum near the upright position. The experimental results show that our RL method can be applied to optimal control problems having continuous state/action spaces and that the method achieves good control with a small number of trial-and-errors.