Scaled free-energy based reinforcement learning for robust and efficient learning in high-dimensional state spaces.

Scaled free-energy based reinforcement learning for robust and efficient learning in high-dimensional state spaces.
复制标题

DOI:
10.3389/fnbot.2013.00003
复制
发表时间:
2013
影响因子:
3.1
通讯作者:
Doya K
Doya K
中科院分区:
计算机科学3区
文献类型:
--
作者:
Elfwing S;Uchibe E;Doya K

文献摘要

参考文献

被引文献

相似文献

基于自由能的强化学习(FERL)被提出用于高维状态和动作空间中的学习,这是标准函数逼近方法无法处理的。在这项研究中,我们提出了基于自由能的强化学习的缩放版本,以实现更稳健、更高效的学习性能。动作值函数由受限玻尔兹曼机的负自由能除以与玻尔兹曼机的大小(本研究中状态节点数量的平方根)相关的常数缩放因子来近似。我们的第一个任务是数字层网格世界任务,其中状态由 MNIST 数据集中的手写数字图像表示。该任务的目的是通过提取隐藏层中与任务相关的特征来研究所提出的方法对相同数字的图像进行聚类以及对与具有相同最佳动作的状态相对应的不同数字的图像进行聚类的能力。我们还测试了该方法对于不同探索计划的鲁棒性,即初始温度和 Softmax 动作选择中温度贴现率的不同设置。我们的第二个任务是机器人视觉导航任务,其中机器人可以通过四个地标下部的不同颜色来学习其位置,并且可以通过地标上部的颜色推断出正确的角球目标区域。状态空间由最多九种不同颜色的二值化相机图像组成,相当于 6642 个二值状态。对于这两个任务,将学习性能与标准 FERL 和函数逼近进行比较,其中动作值函数由两层前馈神经网络逼近。
Free-energy based reinforcement learning (FERL) was proposed for learning in high-dimensional state- and action spaces, which cannot be handled by standard function approximation methods. In this study, we propose a scaled version of free-energy based reinforcement learning to achieve more robust and more efficient learning performance. The action-value function is approximated by the negative free-energy of a restricted Boltzmann machine, divided by a constant scaling factor that is related to the size of the Boltzmann machine (the square root of the number of state nodes in this study). Our first task is a digit floor gridworld task, where the states are represented by images of handwritten digits from the MNIST data set. The purpose of the task is to investigate the proposed method's ability, through the extraction of task-relevant features in the hidden layer, to cluster images of the same digit and to cluster images of different digits that corresponds to states with the same optimal action. We also test the method's robustness with respect to different exploration schedules, i.e., different settings of the initial temperature and the temperature discount rate in softmax action selection. Our second task is a robot visual navigation task, where the robot can learn its position by the different colors of the lower part of four landmarks and it can infer the correct corner goal area by the color of the upper part of the landmarks. The state space consists of binarized camera images with, at most, nine different colors, which is equal to 6642 binary states. For both tasks, the learning performance is compared with standard FERL and with function approximation where the action-value function is approximated by a two-layered feedforward neural network.
DOI: 10.1385/ni:3:3:197
发表时间: 2005-01-01
期刊: NEUROINFORMATICS
影响因子: 3
作者:
Krichmar, JL;Seth, AK;Edelman, GM
通讯作者: Edelman, GM
DOI: 10.1162/089976602753712972
发表时间: 2002-06-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Doya, K;Samejima, K;Kawato, M
通讯作者: Kawato, M
DOI: 10.1073/pnas.0611571104
发表时间: 2007-02-27
影响因子: 11.1
作者:
Fleischer, Jason G.;Gally, Joseph A.;Krichmar, Jeffrey L.
通讯作者: Krichmar, Jeffrey L.
DOI: 10.1177/105971230501300206
发表时间: 2005-06-01
期刊: ADAPTIVE BEHAVIOR
影响因子: 1.6
作者:
Doya, K;Uchibe, E
通讯作者: Uchibe, E
DOI: 10.1162/089976602760128018
发表时间: 2002-08-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Hinton, GE
通讯作者: Hinton, GE