Proposal of Exploitation-Oriented Learning PS-r#
Proposal of Exploitation-Oriented Learning PS-r#
复制标题
DOI:
10.1007/978-3-540-88906-9_1
复制
发表时间:
2008-11
期刊:
影响因子:
--
通讯作者:
K. Miyazaki;S. Kobayashi
中科院分区:
文献类型:
--
作者:
K. Miyazaki;S. Kobayashi
Exploitation-oriented Learning(XoL) is a novel approach to goal-directed learning from interaction. Thoughreinforcement learningis much more focus on the learning and can gurantee the optimality inMarkov Decision Processes(MDPs) environments, XoL aims to learna rational policy, whose expected reward per an action is larger than zero, very quickly. We know PS-r* that is one of the XoL methods. It can learnan useful rational policythat is not inferior to a random walk inPartially Observed Markov Decision Processes(POMDPs) environments where the number of types of a reward is one. However, PS-r* requiresO(MN2) memories whereNandMare the numbers of types of a sensory input and an action.In this paper, we propose PS-r#that can learn an useful rational policy in the POMDPs environments byO(MN) memories. We confirm the effectiveness of PS-r#in numerical examples.