Proposal of Exploitation-Oriented Learning PS-r#

Proposal of Exploitation-Oriented Learning PS-r#
复制标题

DOI:
10.1007/978-3-540-88906-9_1
复制
发表时间:
2008-11
期刊:
--
影响因子:
--
通讯作者:
K. Miyazaki;S. Kobayashi
K. Miyazaki;S. Kobayashi
中科院分区:
其他
文献类型:
--
作者:
K. Miyazaki;S. Kobayashi

文献摘要

相似文献

面向开发的学习(XoL)是从交互中实现目标导向学习的一种新方法。虽然强化学习更注重学习,并且可以保证马尔可夫决策过程(mdp)环境中的最优性,但XoL的目标是非常快速地学习理性策略,其每个动作的预期奖励大于零。我们知道PS-r*是XoL方法之一。它可以在部分可观察马尔可夫决策过程(pomdp)环境中学习有用的理性策略,这种策略不次于随机漫步,其中奖励的类型数量为1。然而,PS-r*需要50 (MN2)个记忆,其中包括感觉输入和动作的类型数量。在本文中,我们提出ps -r#可以通过o (MN)存储器在pomdp环境中学习有用的理性策略。通过数值算例验证了ps -r#的有效性。
Exploitation-oriented Learning(XoL) is a novel approach to goal-directed learning from interaction. Thoughreinforcement learningis much more focus on the learning and can gurantee the optimality inMarkov Decision Processes(MDPs) environments, XoL aims to learna rational policy, whose expected reward per an action is larger than zero, very quickly. We know PS-r* that is one of the XoL methods. It can learnan useful rational policythat is not inferior to a random walk inPartially Observed Markov Decision Processes(POMDPs) environments where the number of types of a reward is one. However, PS-r* requiresO(MN2) memories whereNandMare the numbers of types of a sensory input and an action.In this paper, we propose PS-r#that can learn an useful rational policy in the POMDPs environments byO(MN) memories. We confirm the effectiveness of PS-r#in numerical examples.