Instance-Based Policy Search using Binomial Distribution Crossover and Iterated Refreshment

Instance-Based Policy Search using Binomial Distribution Crossover and Iterated Refreshment
复制标题

使用二项分布交叉和迭代刷新的基于实例的策略搜索

DOI:
10.1109/cec.2006.1688333
复制
发表时间:
2006
期刊:
2006 IEEE International Conference on Evolutionary Computation
影响因子:
--
通讯作者:
S. Kobayashi
S. Kobayashi
中科院分区:
--
文献类型:
--
作者:
Chikao Tsuchiya;Kokolo Ikeda;J. Sakuma;I. Ono;S. Kobayashi

文献摘要

参考文献

被引文献

相似文献

本文描述了一种基于 GA 的强化学习惰性方法。该方法采用数据驱动策略,由实例集和基于实例的操作选择器组成。此功能提供了许多优点。然而,一些困难仍未得到调查。其中之一是巨大而复杂的搜索空间。我们认为保留 GA 群体的特征并引入新特征可以克服这些困难。基于这个思想,我们提出了两种遗传算子;二项式分布交叉 (BDX) 和迭代更新。 BDX生成继承父母特征的后代,并且迭代更新贪婪地引入新特征。由这些算子支持的 GA 被应用于基准任务以展示其能力。各个运营商也从不同的角度进行了调查和讨论。最后,我们为我们的方法提供了优选的参数设置。
This paper describes a GA based lazy approach toward reinforcement learning. This approach employs data-driven policy, which is composed of an instance set and an instance-based action selector. This feature provides a number of advantages. However some difficulties remain uninvestigated. One of them is the huge and complicated search space. We have an idea that preserving characteristics of the GA population and introducing new characteristics can overcome these difficulties. On the basis of this idea, we propose two genetic operators; Binomial Distribution Crossover (BDX) and iterated refreshment. The BDX generates the descendants inheriting the parents' characteristics and the iterated refreshment introduces new characteristics greedily. The GA powered by these operators was applied to the benchmark tasks to demonstrate the ability. Each operator also was investigated and discussed from the various perspectives. Finally, we provide the preferable parameter settings for our method.
DOI: 10.1177/105971239700600201
发表时间: 1997-09
期刊: Adaptive Behavior
影响因子: 1.6
作者:
J. Santamaría;R. Sutton;A. Ram
通讯作者: J. Santamaría;R. Sutton;A. Ram