Behavior coordination for a mobile robot using modular reinforcement learning

Behavior coordination for a mobile robot using modular reinforcement learning
复制标题

使用模块化强化学习的移动机器人的行为协调

DOI:
10.1109/iros.1996.568989
复制
发表时间:
1996
期刊:
Proceedings of IEEE/RSJ International Conference on Intelligent Robots and Systems. IROS '96
影响因子:
--
通讯作者:
K. Hosoda
K. Hosoda
中科院分区:
--
文献类型:
--
作者:
E. Uchibe;M. Asada;K. Hosoda

文献摘要

被引文献

相似文献

通过强化学习方法独立获得的多个行为的协调是将该方法扩展到更大、更复杂的机器人学习任务的问题之一。直接组合各个模块(子任务)的所有状态空间需要大量的学习时间,并且会导致隐藏状态。本文提出了一种模块化学习方法,该方法考虑学习时间和性能之间的权衡来协调多种行为。首先,为了减少学习时间,根据Q学习分别获得的动作值,将整个状态空间分为两类:直接适用其中一种学习行为的区域(不再学习区域),以及由于多个行为的竞争而需要学习的区域(重新学习区域)。其次,根据信息标准,通过模型拟合学习到的动作值来检测隐藏状态。最后,调整重新学习区域的初始动作阀,使其与不再学习区域的值一致。该方法应用于一对一踢足球的机器人。给出了计算机仿真和真实机器人实验,证明了该方法的有效性。
Coordination of multiple behaviors independently obtained by a reinforcement learning method is one of the issues in order for the method to be scaled to larger and more complex robot learning tasks. Direct combination of all the state spaces for individual modules (subtasks) needs enormous learning time, and it causes hidden states. This paper presents a method of modular learning which coordinates multiple behaviors taking account of a trade-off between learning time and performance. First, in order to reduce the learning time the whole state space is classified into two categories based on the action values separately obtained by Q learning: the area where one of the learned behaviors is directly applicable (no more learning area), and the area where learning is necessary due to competition of multiple behaviors (re-learning area). Second, hidden states are detected by model fitting to the learned action values based on the information criterion. Finally, the initial action valves in the re-learning area are adjusted so that they can be consistent with the values in the no more learning area. The method is applied to one to one soccer playing robots. Computer simulation and real robot experiments are given, to show the validity of the proposed method.