课题基金 / 基金详情

Multi-agent Reinforcement Learning Based on Compressed Representation of Decision Policies

Multi-agent Reinforcement Learning Based on Compressed Representation of Decision Policies
基于决策策略压缩表示的多智能体强化学习
批准号:
12680387
负责人:
ONO Norihiko
金额:
$2.3万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2000
资助国家:
日本
项目状态:
已结题
起止时间:
2000 至 2001

项目摘要

项目成果

ONO Norihiko的其他基金

相似基金

相关文献

中文摘要
翻译
据报道,有几种尝试让多个单片强化学习(RL)智能体合成高度协调的行为,以有效地完成它们的共同目标。大多数直接的强化学习应用都不适用于更复杂的多智能体(MA)学习问题,因为每个强化学习智能体的状态空间随着参与联合任务的伙伴智能体的数量呈指数增长。为了弥补多智能体强化学习(MARL)中指数级大的状态空间,我们之前提出了一种模块化方法,并通过应用于多智能体学习问题证明了其有效性。采用模块化方法研究MARL的结果令人鼓舞,但仍存在严重的问题。该方法假设:(i)智能体的所有感官输入和动作输出都是离散值,(ii)所有智能体在规则的时间间隔内完全同步地做出决策,而这种假设在现实世界的多智能体环境中通常不成立。我们提出了另一个MARL框架,该框架可以克服MARL中的状态空间爆炸,该框架基于智能体决策策略的神经网络表示及其用实编码遗传算法的优化,该框架适用于允许单个智能体接收和输出离散/连续值并异步做出决策的多智能体领域。为了证明所提出的框架对实际MARL的有效性,我们将其应用于蜂窝电话系统中的异步多智能体跷跷板平衡问题和动态信道分配问题。结果非常令人鼓舞,而使用任何其他传统MARL框架都无法适当解决这些问题。
英文摘要
Several attempts have been reported to let multiple monolithic reinforcement learning (RL) agents synthesize highly coordinated behavior needed to accomplish their common goal effectively. Most of these straightforward application of RL scale poorly to more complex multi-agent (MA) learning problems, because the state space for each RL agent grows exponentially with the number of its partner agents engaged in the joint task. To remedy the exponentially large state space in multi-agent RL (MARL), we previously proposed a modular approach and demonstrated its effectiveness through the application to the MA learning problems.The results obtained by modular approach to MARL are encouraging, but it still has serious problems. The approach supposes: (i) all the sensory inputs and action outputs for an agent are discrete values, and (ii) all the agents make their decisions totally synchronously at regular time intervals, while such assumption does not hold in real-world multi-agent environments in general.We propose yet another MARL framework which can overcome the state space explosion in MARL, based on neural network representation of the decision policy for an agent and its optimization with a real-coded GA, which is applicable to multi-agent domains where individual agents are allowed to receive and output discrete/continuous values and to make their decisions asynchronously. To show the effectiveness of the proposed framework for real-world MARL, we have applied it to the asynchronous multi-agent seesaw balancing problem and the dynamic channel allocation problem in cellular telephone systems. The results are quite encouraging, while those problems can not be solved appropriately using any other conventional MARL frameworks.
期刊论文(24)
专著(0)
科研奖励(0)
会议论文
Isao Ono, Miyuki Takahashi and Norihiko Ono: "Evolving Neural Networks in Environments with Delayed Rewards by A Real-Coded GA Using Unimodal Normal Distribution Crossover"Proc. 2000 Congress on Evolutionary Computation )CEC2000). 659-666 (2000)
Isao Ono、Miyuki Takahashi 和 Norihiko Ono:“通过使用单峰正态分布交叉的实数编码 GA 在具有延迟奖励的环境中进化神经网络”Proc。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Isao Ono: "Evolving Neural Networks in Environments with Delayed Rewards by A Real-Coded GA using Unimodal Normal Distribution Crossover"Proc.2000 Congress on Evolutionary Computation (CEC2000). 659-666 (2000)
Isao Ono:“使用单峰正态分布交叉通过实数编码 GA 在具有延迟奖励的环境中进化神经网络”Proc.2000 进化计算大会 (CEC2000)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
Isao Ono: "A Genetic Algorithm for Automatically Designing Modular Reinforcement Learning Agents"Proc.2000 Genetic and Evolutionary Conference (GECCO 2000). 203-210 (2000)
Isao Ono:“自动设计模块化强化学习代理的遗传算法”Proc.2000 遗传与进化会议(GECCO 2000)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
共 24 条
    A CO-EVOLUTIONARY MULTI-AGENT REINFORCEMENT LEARNING SCHEME TAKING ACCOUNT OF APPLICATION TO COMPETITIVE ENVIRONMENTS
    • 批准号:
      16500081
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.37万
    • 财政年份:
      2004
    • 负责人:
      ONO Norihiko
    • 依托单位:
    MULTI-AGENT REINFORCEMENT LEARNING WITH NEUROEVOLUTION
    • 批准号:
      14580421
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.3万
    • 财政年份:
      2002
    • 负责人:
      ONO Norihiko
    • 依托单位:
    Synthesis of Coordinated Behavior by Autonomous Agents
    • 批准号:
      10680384
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $0.64万
    • 财政年份:
      1998
    • 负责人:
      ONO Norihiko
    • 依托单位:
    SELF-ORGANIZING MULTI-AGENT SYSTEMS : ARTIFICIAL LIFE APPROACHES
    • 批准号:
      07680402
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $1.47万
    • 财政年份:
      1995
    • 负责人:
      ONO Norihiko
    • 依托单位:
    海外基金