Multiple model-based reinforcement learning

Multiple model-based reinforcement learning
复制标题

DOI:
10.1162/089976602753712972
复制
发表时间:
2002-06-01
期刊:
影响因子:
2.9
通讯作者:
Kawato, M
Kawato, M
中科院分区:
计算机科学4区
文献类型:
--
作者:
Doya, K;Samejima, K;Kawato, M

文献摘要

被引文献

相似文献

我们提出了一个模块化的强化学习架构的非线性,非平稳控制任务,我们称之为多模型为基础的强化学习(MMRL)。其基本思想是根据环境动态的可预测性,将复杂任务分解为空间和时间上的多个域。该系统由多个模块组成,每个模块由状态预测模型和强化学习控制器组成。由预测误差的softmax函数给出的“责任信号”用于对多个模块的输出进行加权,以及门控预测模型和强化学习控制器的学习。我们制定MMRL离散时间,有限状态的情况下,连续时间,连续状态的情况下。MMRL的性能表现为离散的情况下,在一个非平稳狩猎任务在网格世界和连续的情况下,在一个非线性,非平稳的控制任务摆动的钟摆与可变的物理参数。
We propose a modular reinforcement learning architecture for nonlinear, nonstationary control tasks, which we call multiple model-based reinforcement learning (MMRL). The basic idea is to decompose a complex task into multiple domains in space and time based on the predictability of the environmental dynamics. The system is composed of multiple modules, each of which consists of a state prediction model and a reinforcement learning controller. The "responsibility signal," which is given by the softmax function of the prediction errors, is used to weight the outputs of multiple modules, as well as to gate the learning of the prediction models and the reinforcement learning controllers. We formulate MMRL for both discrete-time, finite-state case and continuous-time, continuous-state case. The performance of MMRL was demonstrated for discrete case in a nonstationary hunting task in a grid world and for continuous case in a nonlinear, nonstationary control task of swinging up a pendulum with variable physical parameters.