Strategy Classification in Multi-agent Environment — Applying Reinforcement Learning to Soccer Agents —

Strategy Classification in Multi-agent Environment — Applying Reinforcement Learning to Soccer Agents —
复制标题

多智能体环境中的策略分类——将强化学习应用于足球智能体——

DOI:
10.1109/iros.2004.1389558
复制
发表时间:
2004
期刊:
2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE Cat. No.04CH37566)
影响因子:
--
通讯作者:
K. Hosoda
K. Hosoda
中科院分区:
--
文献类型:
--
作者:
E. Uchibe;M. Asada;K. Hosoda

文献摘要

被引文献

相似文献

本文提出了一种智能体行为分类方法,该方法利用系统识别的方法,通过交互来估计学习者的行为与环境中其他智能体之间的关系。为了识别每个智能体的模型,将赤池信息准则(AIC)应用于典型变量分析(CVA)的结果。接下来,使用基于估计状态向量的强化学习来获得最优行为。将该方法应用于足球机器人。与我们之前的工作不同,该方法可以处理滚动的球。给出了计算机模拟和初步实验,并进行了讨论。构建一个能够学习执行任务的机器人已经被认为是机器人技术和人工智能面临的主要挑战之一。强化学习最近受到越来越多的关注,因为它是一种很少或没有先验知识的机器人学习方法,具有更高的反应性和适应性行为能力(Connel & Mahadevan 1993)。在我们之前的工作(Uchibe, Asada, & Hosoda 1996)中,我们提出了一种模块化强化学习方法,该方法在考虑学习时间和性能之间的权衡的情况下协调多种行为。在多智能体环境中,标准的强化学习算法似乎不适用,因为从学习智能体的角度来看,包括其他学习智能体在内的环境似乎是随机变化的。我们认为在多智能体环境中学习困难主要有两个原因。另一个智能体可能使用随机行动选择器,即使产生相同的感觉,它也可能采取不同的行动。另一个智能体可能有与学习智能体不同的感知(感觉)。这意味着学习代理将无法区分不同的情况,而另一个代理可以做到,反之亦然。因此,除非有明确的沟通,否则即使它的策略是固定的,学习者也不能正确地预测其他代理的行为。对于学习者来说,了解其他智能体的策略并提前预测它们的动作对于成功学习行为至关重要。Littman (Littman 1994)提出了马尔可夫博弈框架,在这个框架中,q学习代理试图在网格世界的零和二人博弈中,学习一种最优的混合策略,以对抗最差的对手。他假设对手的策略被告知学习者(对手试图最小化单个奖励函数,而学习代理则将其最大化)。Sandholm和Crites (Sandholm & Crites 1995)研究了各种Q-learning代理与未知对手进行迭代囚徒困境博弈的能力。他们表明,为了成功地学习,需要足够的先前动作和感觉。Lin (Lin & Mitchell 1992)将基于当前感觉和N个最近的感觉和动作的window-Q与基于循环网络的recurrent- q进行了比较,他发现后者优于前者,因为循环网络可以处理历史特征。然而,一般来说,提前确定神经元的数量和网络的结构仍然是很困难的。如上所述,多智能体环境中的现有方法需要假设其他智能体的策略是固定的,并且为学习者所知,以便学习收敛。因此,需要分类架构来应用强化学习。然而,学习代理所能做的是收集所有观察到的数据,并在观察过程中采取运动命令,并估计被观察代理与学习者行为之间的关系,以便采取适当的行为,尽管由于感知能力的限制,由于部分观察可能无法保证最佳。在本文中,我们提出了一种利用系统识别的方法,通过交互来估计学习者的行为与其他智能体之间的关系。这里,我们把重点放在问题B上,我们假设另一个代理不改变策略。为了识别彼此代理的模型,我们将Akaike的信息准则(AIC) (Akaike 1974)应用于典型变量分析(CVA) (Larimore 1990)的结果,该方法在系统识别领域得到了广泛的应用。我们将提出的方法应用到一个简单的类足球游戏中,其中包括两个主动智能体。agent的任务是区分其他agent的策略。在这里,其他代理包括静止代理(球门和球线)、被动代理(球)和主动代理(对手)。在模型识别后,我们应用强化学习来获取投篮和传球行为。在我们之前的工作中(Asada et al. 1995; Uchibe, Asada, & Hosoda 1996),没有考虑球的大小和位置的变化,因此agent无法在球滚动时获得最佳行为。然而,该方法可以处理移动的球,因为选择了适当的学习状态向量,从而预测了连续的步骤。给出了仿真结果和初步的实际实验,并进行了讨论。Agent Classification Canonical Variate Analysis(CVA)为了学习成功,学习者有必要预测前面提到的后续情况。下面,我们考虑采用一种系统辨识的方法,将电机指令和观测结果分别作为系统的输入和输出。针对多输入多输出(MIMO)组合确定性随机系统的辨识问题,提出了许多算法。与PEM(预测误差法)等“经典”算法相比,子空间系统识别算法(Van Overschee & De Moor 1995)不会受到先验参数化引起的问题的困扰。Larimore的典型变量分析(Canonical Variate Analysis, CVA) (Larimore 1990)就是其中的一种算法,它使用典型相关分析来构造状态估计器。设u(t)∈
This paper proposes a method for agent behavior classification which estimates the relations between the learner’s behaviors and the other agents in the environment through interactions using the method of system identification. In order to identify the model of each agent, Akaike’s Information Criterion(AIC) is applied to the result of Canonical Variate Analysis(CVA). Next, reinforcement learning based on the estimated state vectors is used in order to obtain the optimal behavior. The proposed method is applied to soccer playing robots. Unlike our previous work, the method can cope with a rolling ball. Computer simulations and preliminary experiments are shown and the discussion is given. Introduction Building a robot that learns to perform a task has been acknowledged as one of the major challenges facing Robotics and AI. Reinforcement learning has recently been receiving increased attention as a method for robot learning with little or no a priori knowledge and higher capability of reactive and adaptive behaviors (Connel & Mahadevan 1993). In our previous work (Uchibe, Asada, & Hosoda 1996), we proposed a method of modular reinforcement learning which coordinates multiple behaviors taking account of a trade-off between learning time and performance. In a multi-agent environment, the standard reinforcement learning algorithm does not seem applicable because the environment including the other learning agents seems to change randomly from a viewpoints of the learning agent. We suppose that there are two major reasons why the learning would be difficult in a multi-agent environment. A The other agent may use a stochastic action selector which could take a different action even if the same sensation occurs to it. B The other agent may have a perception (sensation) different from the learning agent’s. This means that the learning agent would not be able to discriminate different situations which the other agent can do, and vice versa. Therefore, the learner cannot predict the other agent behaviors correctly even if its policy is fixed unless explicit communication is available. It is important for the learner to understand the strategies of the other agents and to predict their movements in advance to learn the behaviors successfully. Littman (Littman 1994) proposed the framework of Markov Games in which Q-learning agents try to learn a mixed strategy optimal against the worst possible opponent in a zero-sum 2-player game in a grid world. He assumed that the opponent’s strategy is given to the learner (the opponent tries to minimize a single reward function, while it is to be maximized by the learning agent). Sandholm and Crites (Sandholm & Crites 1995) studied the ability of a variety of Q-learning agents to play iterated prisoner’s dilemma game against an unknown opponent. They showed that adequate previous moves and sensations are needed in order to learn successfully. Lin (Lin & Mitchell 1992) compared window-Q based on both the current sensation and the N most recent sensations and actions with recurrent-Q based on a recurrent network, and he showed the latter is superior to the former because a recurrent network can cope with historical features. However, generally speaking it is still difficult to determine the number of neurons and the structures of network in advance. As described above, existing methods in multi agent environments need the assumption that the policy of other agent is fixed and known to the learner in order for the learning to converge. Therefore, the classification architecture is required to apply the reinforcement learning. However, what the learning agent can do is to collect all the observed data with motor commands taken during the observation and to estimate the relationship between the observed agents and the learner’s behaviors in order to take an adequate behavior although it might not be guaranteed as optimal because of partial observation due to the limitation of sensing capability. In this paper, we propose a method which estimates the relations between the learner’s behaviors and the other agents through interactions using the method of system identification. Here, we put our emphasis on the problem B, and we assume that the other agent does not change the strategy. In order to identify the model of each other agent, we apply Akaike’s Information Criterion(AIC) (Akaike 1974) to the result of Canonical Variate Analysis(CVA) (Larimore 1990), which is widely used in the field of system identification. We apply the proposed method to a simple soccerlike game including two active agents. The task of the agent is to discriminate the strategy of the other agents. Here, the other agents consist of the stationary agent (the goal and the line), passive agent (the ball) and active agent (the opponent). After the model identification, we apply reinforcement learning in order to acquire shooting and passing behaviors. In our previous work (Asada et al. 1995; Uchibe, Asada, & Hosoda 1996), the changes in size and position of the ball are not considered, therefore the agent could not acquire optimal behavior when the ball is rolling. However, the proposed method can cope with the moving ball because state vector for learning is selected appropriately so as to predict the successive steps. Simulation results and preliminary real experiments are shown and the discussion is given. Agent Classification Canonical Variate Analysis(CVA) In order to succeed in learning, it is necessary for the learner to predict the subsequent situations as mentioned above. In the following, we consider to utilize a method of system identification, regarding motor command and observation results as the input and the output of the system respectively. A number of algorithms to identify multiinput multi-output (MIMO) combined deterministicstochastic systems have been proposed. In contrast to ‘classical’ algorithms such as PEM (Prediction Error Method), the subspace system identification algorithms (Van Overschee & De Moor 1995) do not suffer from the problems caused by a priori parameterizations. Larimore’s Canonical Variate Analysis (CVA) (Larimore 1990) is one of such algorithms, which uses canonical correlation analysis to construct a state estimator. Let u(t) ∈