Delayed reward-based genetic algorithms for partially observable Markov decision problems

Delayed reward-based genetic algorithms for partially observable Markov decision problems
复制标题

针对部分可观察马尔可夫决策问题的基于延迟奖励的遗传算法

DOI:
10.1002/scj.10230
复制
发表时间:
2004
期刊:
Syst. Comput. Jpn.
影响因子:
--
通讯作者:
H. Takeda
H. Takeda
中科院分区:
--
文献类型:
--
作者:
Yoshihide Yamashiro;A. Ueno;H. Takeda

文献摘要

被引文献

相似文献

强化学习通常涉及假设马尔可夫特征。然而,agent不能总是完全观察环境,在这种情况下,不同的状态被观察为相同的状态。在本研究中,作者开发了一种基于延迟奖励的POMDP遗传算法(DRGA),作为解决具有感知混叠问题的部分可观察马尔可夫决策问题(POMDP)的手段。DRGA将POMDP分解为几个子任务,然后通过将代理分解为几个子代理来解决POMDP问题。每个子代理根据环境的延迟奖励获取适应环境的策略,这些策略使用基于延迟奖励的遗传算法进化。智能体通过结合自然选择后留下的有效策略来适应环境。作者将该方法应用于感知有限的迷宫搜索问题,以证明其有效性。©2004 Wiley期刊公司系统比较,35(2):66 - 78,2004;在线发表于Wiley InterScience (www.interscience)。wiley.com)。DOI 10.1002 / scj.10230
Reinforcement learning often involves assuming Markov characteristics. However, the agent cannot always observe the environment completely, and in such cases, different states are observed as the same state. In this research, the authors develop a Delayed Reward-based Genetic Algorithm for POMDP (DRGA) as a means to solve a partially observable Markov decision problem (POMDP) which has such perceptual aliasing problems. The DRGA breaks down the POMDP into several subtasks, and then solves the POMDP by breaking down the agent into several subagents. Each subagent acquires policies adapted to the environment based on the delayed rewards from the environment, and these policies are evolved using a genetic algorithm based on the delayed rewards. The agent adapts to the environment by combining effective policies that remain after natural selection. The authors apply this method to maze search problems in which perception is limited in order to demonstrate its validity. © 2004 Wiley Periodicals, Inc. Syst Comp Jpn, 35(2): 66–78, 2004; Published online in Wiley InterScience (www.interscience. wiley.com). DOI 10.1002/scj.10230