Cooperative behavior acquisition by asynchronous policy renewal that enables simultaneous learning in multiagent environment

Cooperative behavior acquisition by asynchronous policy renewal that enables simultaneous learning in multiagent environment
复制标题

通过异步策略更新获取合作行为,实现多智能体环境中的同步学习

DOI:
10.1109/irds.2002.1041682
复制
发表时间:
2002
期刊:
IEEE/RSJ International Conference on Intelligent Robots and Systems
影响因子:
--
通讯作者:
K. Hosoda
K. Hosoda
中科院分区:
--
文献类型:
--
作者:
Shoichi Ikenoue;M. Asada;K. Hosoda

文献摘要

被引文献

相似文献

提出了一种多智能体环境下的同步学习方法,以促进协作行为。每个agent有一个策略和一个动作值函数,前者是基于前一阶段更新的动作值函数来执行动作,后者是基于当前策略所经历的事件进行学习。这使得所有智能体的行为都基于固定的策略,因此除了依赖于每个智能体学习进度的更新周期外,可以避免非马尔可夫问题。为了避免由于动作值函数的异步更新而产生局部最大值,首先给出了乐观的动作值,避免了探索过程陷入局部最大值。实验结果应用于动态多智能体环境下的一个协作任务RoboCup,并进行了讨论。
This paper presents a method for simultaneous learning in multiagent environment to facilitate cooperative behavior. Each agent has one policy and one action value function: the former is for action execution based on the action value function updated in the previous stage, and the latter is for learning based on the episodes experienced by the current policy. This makes all agents behave based on the fixed policies, so that the non-Markovian problem can be avoided except for the update periods that depend on the learning progress of each agent. In order to avoid the local maxima due to such asynchronous renewal of action value functions, optimistic action values are given initially, which helps to avoid the exploration process being trapped in local maxima. The experimental results applied to one of the cooperative tasks in a dynamic, multiagent environment, RoboCup, is shown and a discussion is given.