MULTI-AGENT REINFORCEMENT LEARNING WITH NEUROEVOLUTION
MULTI-AGENT REINFORCEMENT LEARNING WITH NEUROEVOLUTION
批准号:
14580421
负责人:
ONO Norihiko
金额:
$2.3万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2002
资助国家:
日本
项目状态:
已结题
起止时间:
2002 至 2003
中文摘要
据报道,有几种尝试让多个单片强化学习(RL)智能体合成高度协调的行为,以有效地完成它们的共同目标。大多数直接的强化学习应用都不适用于更复杂的多智能体学习问题,因为每个强化学习智能体的状态空间随着参与联合任务的伙伴智能体的数量呈指数增长。为了应对多智能体强化学习(MARL)中指数级大的状态空间,我们提出了一种基于智能体决策策略的神经网络表示及其实数编码遗传算法优化的MARL方案,并将该方案应用于其他传统MARL方案无法解决的多智能体学习问题,表明了该方案的有效性。例如蜂窝电话系统中的异步多智能体跷跷板平衡问题和动态信道分配问题。然而,由于该方案需要大量的计算资源,我们不能将其直接应用于大规模多智能体系统(MASs)的设计问题。为了弥补这一缺陷,我们提出了一种大规模MAS的分层设计方案,该方案将MAS的整个任务分层分解为子任务,并利用上述MARL方案对每个子任务进行优化。将该设计方案应用于机器人世界杯足球队设计问题,将足球队任务分解为足球agent的原始动作、agent之间的交互和agent之间的协调,证明了该设计方案的有效性。
英文摘要
Several attempts have been reported to let multiple monolithic reinforcement learning (RL) agents synthesize highly coordinated behavior needed to accomplish their common goal effectively. Most of these straightforward application of RL scale poorly to more complex multi-agent learning problems, because the state space for each RL agent grows exponentially with the number of its partner agents engaged in the joint task.To cope with the exponentially large state space in multi-agent RL (MARL), we previously proposed a MARL scheme, based on neural network representation of the decision policy for an agent and its optimization with a real-coded GA, and showed the effectiveness of the scheme through its application to those multi-agent learning problems that can not be solved appropriately using any other conventional MARL scheme, such as the asynchronous multi-agent seesaw balancing problem and the dynamic channel allocation problem in cellular telephone systems.However, we can not apply the scheme directly to the design problems of large-scale multi-agent systems (MASs), because the scheme needs a huge amount of computation resources. To remedy the drawback, we propose a hierarchical design scheme of a large-scale MAS, which simply decomposes the whole task of the MAS into its subtasks hierarchically and optimizes each of the subtasks with the above-mentioned MARL scheme. The effectiveness of the design scheme is shown through its application to the RoboCup soccer team design problem where the task of the team is decomposed into the primitive actions by the soccer agents, interaction among the actions and coordination by the agents.
期刊论文(47)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
阿部 哲: "リカレントニューラルネットの構造と重みの進化的最適化のための世代交代モデルの提案"第47回システム制御情報学会研究発表講演会講演論文集. (2003)
Satoshi Abe:“针对循环神经网络结构和权重的进化优化的一代交替模型的提议”第 47 届系统、控制和信息工程师学会年会论文集(2003 年)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
"Gridifying" A Parallel NSS-EA using The Improved GOGA Framework And Its Performance Evaluation on OBI Grid
使用改进的GOGA框架“网格化”并行NSS-EA及其在OBI网格上的性能评估
DOI:
--
发表时间:
2004
期刊:
First International Workshop on Life Science Grid (LSGRID2004)
影响因子:
--
作者:
[Hiroaki Imade, Naoaki Mizuguchi, Isao Ono, Norihiko Ono, Masahiro Okamoto]
通讯作者:
Masahiro Okamoto
An Evolutionary Algorithm Taking Account of Mutual Interactions among Substances for Inference of Genetic Networks
一种考虑物质间相互作用的遗传网络推理进化算法
DOI:
--
发表时间:
2004
期刊:
Proc.the 2004 Congress on Evolutionary Computation (CEC2004)
影响因子:
--
作者:
[Isao Ono, Yoshiaki Seike, Ryohei Morishita, Norihiko Ono, Masahiko Nakatsui, Masahiro Okamoto]
通讯作者:
Masahiro Okamoto
橋 勇人: "対戦型ゲーム戦略の共進化的獲得に関する実験的考察"第48回システム制御情報学会研究発表講演会講演論文集. (2004)
Hayato Hashi:“竞争性游戏策略共同进化习得的实验研究”第 48 届系统、控制和信息工程师协会年会论文集(2004 年)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
A Framework of Grid-Oriented Genetic Algorithms for Large-Scale Optimization in Bioinformatics
用于生物信息学大规模优化的面向网格遗传算法框架
DOI:
--
发表时间:
2003
期刊:
Proc.the 2003 Congress on Evolutionary Computation (CEC2003)
影响因子:
--
作者:
[Hiroaki Imade, Ryohei Morishita, Isao Ono, Norihiko Ono, Masahiro Okamoto]
通讯作者:
Masahiro Okamoto
共 30 条
A CO-EVOLUTIONARY MULTI-AGENT REINFORCEMENT LEARNING SCHEME TAKING ACCOUNT OF APPLICATION TO COMPETITIVE ENVIRONMENTS
-
批准号:16500081
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.37万
-
财政年份:2004
-
负责人:ONO Norihiko
-
依托单位:
Multi-agent Reinforcement Learning Based on Compressed Representation of Decision Policies
-
批准号:12680387
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.3万
-
财政年份:2000
-
负责人:ONO Norihiko
-
依托单位:
Synthesis of Coordinated Behavior by Autonomous Agents
-
批准号:10680384
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$0.64万
-
财政年份:1998
-
负责人:ONO Norihiko
-
依托单位:
SELF-ORGANIZING MULTI-AGENT SYSTEMS : ARTIFICIAL LIFE APPROACHES
-
批准号:07680402
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.47万
-
财政年份:1995
-
负责人:ONO Norihiko
-
依托单位:
海外基金