EM-based policy hyper parameter exploration: application to standing and balancing of a two-wheeled smartphone robot

EM-based policy hyper parameter exploration: application to standing and balancing of a two-wheeled smartphone robot
复制标题

基于EM的策略超参数探索:应用于两轮智能手机机器人的站立和平衡

DOI:
10.1007/s10015-015-0260-7
复制
发表时间:
2016
影响因子:
0.9
通讯作者:
K. Doya
K. Doya
中科院分区:
--
文献类型:
--
作者:
Jiexin Wang;E. Uchibe;K. Doya

文献摘要

被引文献

相似文献

本文提出了一种新的策略搜索算法,称为基于em的策略超参数探索(EPHE),该算法集成了两种强化学习算法:策略梯度与参数探索(PGPE)和基于em的奖励加权回归。与PGPE一样,EPHE使用从策略超参数(均值和方差)给出的先验分布中采样的策略参数来评估每个事件中的确定性策略。基于em的奖励加权回归,通过奖励加权平均来更新策略超参数,从而不需要梯度计算和学习率的调整。在摆起任务、推车杆平衡任务和两轮智能手机机器人站立与平衡仿真的基准测试中对所提方法进行了验证。实验结果表明,即使对于具有不连续的任务,EPHE也可以在不调整学习率的情况下实现高效的学习。
This paper proposes a novel policy search algorithm called EM-based Policy Hyper Parameter Exploration (EPHE) which integrates two reinforcement learning algorithms: Policy Gradient with Parameter Exploration (PGPE) and EM-based Reward-Weighted Regression. Like PGPE, EPHE evaluates a deterministic policy in each episode with the policy parameters sampled from a prior distribution given by the policy hyper parameters (mean and variance). Based on EM-based Reward-Weighted Regression, the policy hyper parameters are updated by reward-weighted averaging so that gradient calculation and tuning of the learning rate are not required. The proposed method is tested in the benchmarks of pendulum swing-up task, cart-pole balancing task and simulation of standing and balancing of a two-wheeled smartphone robot. Experimental results show that EPHE can achieve efficient learning without learning rate tuning even for a task with discontinuities.