Learning Coordinated Behavior in a Continuous Environment

Learning Coordinated Behavior in a Continuous Environment
复制标题

在连续环境中学习协调行为

DOI:
10.1007/3-540-62934-3_42
复制
发表时间:
1996
期刊:
--
影响因子:
--
通讯作者:
Yoshihiro Fukuta
Yoshihiro Fukuta
中科院分区:
--
文献类型:
--
作者:
N. Ono;Yoshihiro Fukuta

文献摘要

被引文献

相似文献

人们已经做出了有趣的努力,让多个代理学习如何使用各种强化学习算法进行适当的交互。然而,在大多数情况下,假定每个代理的状态空间是离散的。目前还不清楚多个强化学习代理如何有效地在连续状态空间中获得适当的协调行为。本研究的目的是探索q学习在多智能体连续环境中的潜在适用性,当与基于CMAC的泛化技术结合使用时。我们考虑了多智能体块推送问题的改进版本,其中两个学习智能体在连续环境中相互作用以实现它们的共同目标。为了允许我们的智能体处理二维向量值输入,我们应用了基于cmac的q -学习算法。这是l - j . lin的sqconalgorithm的一个变体。目标是逐步细化一组cmac,这些cmac可以近似地为学习智能体提供最优策略下的动作值函数。通过模拟运行,我们对基于块推送cmac的q学习代理的性能进行了定量和定性评估。虽然它并不打算模拟任何特定的现实世界问题,但结果令人鼓舞。
Interesting efforts have been made to let multiple agents learn to appropriately interact, using various reinforcement-learning algorithms. In most of these cases, however, the state space for each agent is supposed discrete. It is not clear how effectively multiple reinforcementlearning agents are able to acquire appropriate coordinated behavior in continuous state spaces. The objective of this research is to explore the potential applicability of Q-learning in multi-agent continuous environments, when applied in conjunction with a generalization technique based on CMAC. We consider a modified version of the multi-agent block pushing problem, where two learning agents are interacting in a continuous environment to accomplish their common goal. To allow our agent to treat two-dimensional vector-valued inputs, we applied a CMAC-based Q-learning algorithm. This is a variant of L.-J.Lin'sQCONalgorithm. The objective is to incrementally elaborate a set of CMACs which can approximately provide the action value function under an optimal policy for the learning agent. The performance of our block pushing CMAC-based Q-learning agents is evaluated quantitatively and qualitatively through simulation runs. Although it is not intended to model any particular real world problem, the results are encouraging.