Permutation based task transfer for genetic programming
Permutation based task transfer for genetic programming
批准号:
RGPIN-2015-06117
负责人:
Heywood, Malcolm
金额:
$1.31万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2018
资助国家:
加拿大
项目状态:
已结题
起止时间:
2018-01-01 至 2019-12-31
中文摘要
本研究建议的一般背景是遗传编程(GP),应用于学习决策策略的代理在延迟支付环境(或强化学习)。该提案的具体重点在于制定一个框架,以便有系统地将GP扩展到比以前所考虑的更困难的延迟付款的任务版本。特别是,我们感兴趣的情景是,一个较简单的初始“源”任务的解决方案随后被“转移”到一个更困难但相关的(目标)任务;或一种形式的迁移学习。这项工作的洞察力是利用GP的能力来识别利用状态变量子集的解决方案。然后,这为重新部署在源任务下发现的解决方案提供了基础,以便可以解决更困难的任务。采用这种方法的潜在好处是:1)不需要从头开始连续地重新发现策略;2)在最终目标任务的解决方案中增加了成功/或更好的解决方案;以及3)相对于为每个任务寻找解决方案而言,降低了计算开销。*将使用两个特定的目标域来说明该方法:1)学习在RoboCup Keepaway的连续值模拟2D世界中踢足球的策略;2)学习用于求解3x3Rubik立方体的一般策略。Keepway足球任务代表了多智能体学习的基准,因此有以前的结果以及提出难度递增的任务的历史。魔方任务以前几乎没有作为学习任何形式的算法的基准的历史。相反,解决方案采取了部署某种形式的全面搜索的形式。这两个任务都是复杂任务域的例子,它们具有非常大的状态空间(合法状态的潜在数量),但具有GP应该能够发现的特定实例的基本属性(规则)。这项研究的基本假设是,一旦确定了用于解决初始任务的某个子集的上下文相关策略的实例,那么我们应该能够将其作为基础,通过策略对状态变量的引用的确定性变化过程,将其推广到更多的任务实例。当然,为了扩大搜索过程,我们需要避免人为地将病理引入搜索。简而言之,用于构建后续策略的初始策略的总和需要超过其各部分的总和。*这些目标的成功将为将GP扩展到具有延迟回报的广泛任务提供一个总体框架。这类任务受到GP社区的广泛关注,因为它们代表了一些最昂贵的(如果不是最昂贵的)应用GP的任务域集合。此外,人们普遍认为这两个任务领域特别具有挑战性。
英文摘要
The general context for this research proposal is that of genetic programming (GP) as applied to learning decision making policies for agents operating in environments with delayed payoff (or reinforcement learning). The specific focus of the proposal lies in developing a framework for systematically scaling GP to more difficult versions of tasks with delayed payoff than have previously been considered. In particular we are interested in scenarios in which solutions for a simpler initial `source' task are then `transferred' to a more difficult but related (target) task; or a form of transfer learning. The insight of this work is to make use of the capability of GP to identify solutions that make use of subsets of state variables. This then provides the basis for redeploying solutions discovered under the source task such that more difficult tasks can be solved. The potential benefits of adopting such an approach are that: 1) it is not necessary to continuously rediscover policies from scratch; 2) increased success in / or better solutions to the ultimate target task; and 3) lower computational overhead as measured against finding solutions to each task.******Two specific target domains will be used to illustrate the approach: 1) learning policies to play soccer in the continuous valued simulated 2D world of RoboCup keepaway; 2) learning general policies for solving the 3 by 3 Rubik cube. The keepway soccer task represents a benchmark for multi-agent learning, hence has a history of previous results as well as posing tasks of incrementally increasing difficulty. The Rubik cube task has had little previous history as a benchmark for learning algorithms of any form. Instead solutions have taken the form of deploying some form of exhaustive search. Both tasks represent examples of complex task domains that have very large state-spaces (potential number of legal states), but possess underlying properties (regularities) that GP should be able to discover specific instances of. The basic hypothesis of this research is that once an instance of a context dependent strategy is identified for solving some subset of an initial task, then we should be able to use this as the basis for generalizing to many more instances of the task through a deterministic process of variation in the policy's references to the state variables. Naturally, for the process to scale, we need to avoid artificially introducing pathologies into the search. In short, the sum of initial policies from which a later policy is constructed needs to exceed the mere sum of its parts.***Success in these objectives would provide a general framework for scaling GP to a wide range of tasks with delayed payoff. Such tasks are of widespread interest to the GP community because they represent some of the most expensive, if not the most expensive, set of task domains for applying GP to. Moreover, the two task domains are widely acknowledged to be of a particularly challenging nature.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Scaling Genetic Programming to Complex Reinforcement Learning Tasks
-
批准号:RGPIN-2020-04438
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2022
-
负责人:Heywood, Malcolm
-
依托单位:
Scaling Genetic Programming to Complex Reinforcement Learning Tasks
-
批准号:RGPIN-2020-04438
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2021
-
负责人:Heywood, Malcolm
-
依托单位:
Scaling Genetic Programming to Complex Reinforcement Learning Tasks
-
批准号:RGPIN-2020-04438
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2020
-
负责人:Heywood, Malcolm
-
依托单位:
Permutation based task transfer for genetic programming
-
批准号:RGPIN-2015-06117
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2019
-
负责人:Heywood, Malcolm
-
依托单位:
Coevolutionary automatic game content generation of physics and flighting style games
-
批准号:499792-2016
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$5.34万
-
财政年份:2018
-
负责人:Heywood, Malcolm
-
依托单位:
Permutation based task transfer for genetic programming
-
批准号:RGPIN-2015-06117
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2017
-
负责人:Heywood, Malcolm
-
依托单位:
Coevolutionary automatic game content generation of physics and flighting style games
-
批准号:499792-2016
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$5.34万
-
财政年份:2017
-
负责人:Heywood, Malcolm
-
依托单位:
Permutation based task transfer for genetic programming
-
批准号:RGPIN-2015-06117
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2016
-
负责人:Heywood, Malcolm
-
依托单位:
Coevolutionary automatic game content generation of physics and flighting style games
-
批准号:499792-2016
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$5.34万
-
财政年份:2016
-
负责人:Heywood, Malcolm
-
依托单位:
Constructing risk predictors for mobile device behaviour analytics
-
批准号:485070-2015
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2015
-
负责人:Heywood, Malcolm
-
依托单位:
Evolving under tasks of incomplete information: streaming and self play
-
批准号:451239-2013
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$6.41万
-
财政年份:2015
-
负责人:Heywood, Malcolm
-
依托单位:
Permutation based task transfer for genetic programming
-
批准号:RGPIN-2015-06117
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2015
-
负责人:Heywood, Malcolm
-
依托单位:
Game Server Network Analysis Engine
-
批准号:489094-2015
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2015
-
负责人:Heywood, Malcolm
-
依托单位:
EEG artifact removal under minimal sensor redundancy
-
批准号:471475-2014
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2014
-
负责人:Heywood, Malcolm
-
依托单位:
Continuous symbiotic program evolution
-
批准号:238791-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2014
-
负责人:Heywood, Malcolm
-
依托单位:
Evolving under tasks of incomplete information: streaming and self play
-
批准号:451239-2013
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$6.41万
-
财政年份:2014
-
负责人:Heywood, Malcolm
-
依托单位:
Continuous symbiotic program evolution
-
批准号:238791-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2013
-
负责人:Heywood, Malcolm
-
依托单位:
Evolving under tasks of incomplete information: streaming and self play
-
批准号:451239-2013
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$6.41万
-
财政年份:2013
-
负责人:Heywood, Malcolm
-
依托单位:
Continuous symbiotic program evolution
-
批准号:238791-2010
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.82万
-
财政年份:2012
-
负责人:Heywood, Malcolm
-
依托单位:
Pattern Validation in Video Lottery Gaming
-
批准号:408123-2010
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$6.81万
-
财政年份:2011
-
负责人:Heywood, Malcolm
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
Incentive and governance schenism study of corporate green washing behavior in China: Based on an integiated view of econfiguration of environmental authority and decoupling logic
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:YU BYUNGJUN
-
依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
-
批准号:W2433169
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:HAOFEI ZHANG
-
依托单位:
含Re、Ru先进镍基单晶高温合金中TCP相成核—生长机理的原位动态研究
-
批准号:52301178
-
项目类别:青年科学基金项目
-
资助金额:30.00万元
-
批准年份:2023
-
负责人:夏万顺
-
依托单位:
NbZrTi基多主元合金中化学不均匀性对辐照行为的影响研究
-
批准号:12305290
-
项目类别:青年科学基金项目
-
资助金额:30.00万元
-
批准年份:2023
-
负责人:苏钲雄
-
依托单位:
眼表菌群影响糖尿病患者干眼发生的人群流行病学研究
-
批准号:82371110
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:邹海东
-
依托单位:
镍基UNS N10003合金辐照位错环演化机制及其对力学性能的影响研究
-
批准号:12375280
-
项目类别:面上项目
-
资助金额:53.00万元
-
批准年份:2023
-
负责人:黄鹤飞
-
依托单位:
CuAgSe基热电材料的结构特性与构效关系研究
-
批准号:22375214
-
项目类别:面上项目
-
资助金额:50.00万元
-
批准年份:2023
-
负责人:周钲洋
-
依托单位:
A study on prototype flexible multifunctional graphene foam-based sensing grid (柔性多功能石墨烯泡沫传感网格原型研究)
-
批准号:--
-
项目类别:--
-
资助金额:20万元
-
批准年份:2020
-
负责人:SAGAR RIZWAN UR REHMAN
-
依托单位:
基于大数据定量研究城市化对中国季节性流感传播的影响及其机理
-
批准号:82003509
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:雷浩
-
依托单位: