Multi-agent reinforcement learning: using macro actions to learn a mating task

Multi-agent reinforcement learning: using macro actions to learn a mating task
复制标题

多智能体强化学习:使用宏观动作来学习交配任务

DOI:
10.1109/iros.2004.1389904
复制
发表时间:
2004
期刊:
2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE Cat. No.04CH37566)
影响因子:
--
通讯作者:
Henrik I. Christensen
Henrik I. Christensen
中科院分区:
--
文献类型:
--
作者:
Stefan Elfwing;E. Uchibe;K. Doya;Henrik I. Christensen

文献摘要

被引文献

相似文献

标准的强化学习方法是低效的,往往不足以学习合作的多智能体任务。对于这些类型的任务,一个代理的行为强烈依赖于与其他代理的动态交互,而不仅仅是与标准强化学习中的静态环境的交互。因此,学习的成功与智能体预测其他智能体行为的能力相关联。在这项研究中,我们试图通过添加一些简单的宏动作来克服这个问题,这些动作在时间上延长了一个以上的时间步。宏动作通过使状态空间的搜索更有效来改进学习,从而使其他代理的行为更可预测。在这项研究中,我们已经考虑了合作交配任务,这是我们的目标,以执行体现进化的第一步,进化选择过程是任务的一个组成部分。我们表明,在模拟和硬件中,在没有宏动作的学习的情况下,代理无法学习有意义的行为。相比之下,对于宏动作的学习,智能体在合理的时间内学习到良好的交配行为,无论是在模拟还是硬件中。
Standard reinforcement learning methods are inefficient and often inadequate for learning cooperative multi-agent tasks. For these kinds of tasks the behavior of one agent strongly depends on dynamic interaction with other agents, not only with the interaction with a static environment as in standard reinforcement learning. The success of the learning is therefore coupled to the agents' ability to predict the other agents behaviors. In this study we try to overcome this problem by adding a few simple macro actions, actions that are extended in time for more than one time step. The macro actions improve the learning by making search of the state space more effective and thereby making the behavior more predictable for the other agent. In this study we have considered a cooperative mating task, which is the first step towards our aim to perform embodied evolution, where the evolutionary selection process is an integrated part of the task. We show, in simulation and hardware, that in the case of learning without macro actions, the agents fail to learn a meaningful behavior. In contrast, for the learning with macro action the agents learn a good mating behavior in reasonable time, in both simulation and hardware.