Totally model-free reinforcement learning by actor-critic Elman networks in non-Markovian domains

Totally model-free reinforcement learning by actor-critic Elman networks in non-Markovian domains
复制标题

非马尔可夫领域中演员批评家 Elman 网络的完全无模型强化学习

DOI:
10.1109/ijcnn.1998.687169
复制
发表时间:
1998
期刊:
1998 IEEE International Joint Conference on Neural Networks Proceedings. IEEE World Congress on Computational Intelligence (Cat. No.98CH36227)
影响因子:
--
通讯作者:
Stuart E. Dreyfus
Stuart E. Dreyfus
中科院分区:
--
文献类型:
--
作者:
Eiji Mizutani;Stuart E. Dreyfus

文献摘要

被引文献

相似文献

我们描述了一个非马尔可夫域中的行动者-评论家强化学习代理如何以完全无模型的方式找到最优的行动序列;也就是说,代理既不学习转移概率和相关的奖励,也不学习状态空间应该增加多少,以便马尔可夫属性保持不变。特别是,我们采用Elman型递归神经网络来解决非马尔可夫问题,因为Elman型网络能够隐式地自动呈现马尔可夫过程。一个标准的“演员-评论家”神经网络模型有两个独立的组成部分:动作(演员)网络和价值(评论家)网络。然而,在动物的大脑中,这两个可能并不明显,而是以某种方式交织在一起。因此,我们构建了一个Elman网络与两个输出节点:演员节点和评论家节点,和共享的隐藏层的一部分被反馈作为上下文层,其功能作为一个历史记忆,以产生敏感性非马尔可夫依赖。智能体探索小规模的三阶段和四阶段三角形路径网络,以学习最佳的行动序列,最大化与其从顶点到顶点的过渡相关的总价值(或奖励)。所提出的问题有确定性的过渡和奖励与每个允许的行动(虽然可以是随机的),并呈现非马尔可夫的奖励依赖于较早的过渡。由于神经模型自由学习的性质,即使在小规模的路径问题中,智能体也需要多次迭代才能找到最佳动作。
We describe how an actor-critic reinforcement learning agent in a non-Markovian domain finds an optimal sequence of actions in a totally model-free fashion; that is, the agent neither learns transitional probabilities and associated rewards, nor by how much the state space should be augmented so that the Markov property holds. In particular, we employ an Elman-type recurrent neural network to solve non-Markovian problems since an Elman-type network is able to implicitly and automatically render the process Markovian. A standard "actor-critic" neural network model has two separate components: the action (actor) network and the value (critic) network. In animal brains, however, those two presumably may not be distinct, but rather somehow entwined. We thus construct one Elman network with two output nodes: actor node and critic node, and a portion of the shared hidden layer is fed back as the context layer, which functions as a history memory to produce sensitivity to non-Markovian dependencies. The agent explores small-scale three and four-stage triangular path-networks to learn an optimal sequence of actions that maximizes total value (or reward) associated with its transition from vertex to vertex. The posed problem has deterministic transition and reward associated with each allowable action (although either could be stochastic) and is rendered non-Markovian by the reward being dependent on an earlier transition. Due to the nature of neural model-free learning, the agent needs many iterations to find the optimal actions even in small-scale path problems.