Experiment of reinforcement learning with extremum seeking

Experiment of reinforcement learning with extremum seeking
复制标题

DOI:
10.1109/ict-ispc.2017.8075301
复制
发表时间:
2017-05
期刊:
2017 6th ICT International Student Project Conference (ICT-ISPC)
影响因子:
--
通讯作者:
Megumi Miyashita;Ryo Hirotani;S. Yano;T. Kondo
Megumi Miyashita;Ryo Hirotani;S. Yano;T. Kondo
中科院分区:
其他
文献类型:
--
作者:
Megumi Miyashita;Ryo Hirotani;S. Yano;T. Kondo

文献摘要

相似文献

最近的研究关注黑盒优化问题和强化学习(RL)问题之间的相似性。他们期望黑盒优化算法能够解决强化学习问题。在黑盒优化算法中,极值搜索(ES)是一种值得注意的算法。过去的两项研究隐式地使用 ES 解决了 RL 问题,但它们的问题设置仅限于线性确定性系统。在本研究中,我们提出了一种新颖的算法来使用 ES 解决更一般的 RL 问题。具体来说,我们采用新的目标函数和在线优化技术来解决随机非线性状态转换环境的问题。作为实验,所提出的方法解决了两个任务。首先,它解决了一个机器人手臂任务,以便与之前的方法 PoWER 进行比较。其次,它解决了一维到达任务以了解探索方式的效果。结果,我们发现我们的算法在适应性方面表现出比 PoWER 更好的性能,并阐明了初始探索噪声强烈影响搜索方向。
Recent studies pay attention to the similarity between black-box optimization problem and reinforcement learning (RL) problem. They expect that black-box optimization algorithm can solve RL problems. Among the black-box optimization algorithms, extremum seeking (ES) is a notable one. Two past studies implicitly solved RL problem using ES, but their problem settings were limited to linear deterministic system. In this study, we propose the novel algorithm to solve a more general RL problem using ES. Specifically, we employ a new objective function and online optimization technique to solve a problem with stochastic non-linear state transition environment. As experiment, the proposed method solves two tasks. First, it solves a robot arm task in order to compare with the previous method PoWER. Second, it solves the one-dimensional reaching task to know the effects of exploration manner. As a result, we found that our algorithm showed better performance in adaptability than PoWER, and clarified that the initial exploration noise strongly affects the direction of search.