Coordination of multiple behaviors acquired by a vision-based reinforcement learning

Coordination of multiple behaviors acquired by a vision-based reinforcement learning
复制标题

基于视觉的强化学习获得的多种行为的协调

DOI:
10.1109/iros.1994.407484
复制
发表时间:
1994
期刊:
Proceedings of IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS'94)
影响因子:
--
通讯作者:
K. Hosoda
K. Hosoda
中科院分区:
--
文献类型:
--
作者:
M. Asada;E. Uchibe;S. Noda;Sukoya Tawaratsumida;K. Hosoda

文献摘要

被引文献

相似文献

提出了一种通过协调基于视觉的强化学习获得的多个行为来完成由多个子任务组成的整体任务的方法。首先,实现相应子任务的个体行为通过Q-learning独立获得,Q-learning是一种广泛使用的强化学习方法。每个学习到的行为都可以用环境状态和机器人动作的动作值函数来表示。其次,考虑了三种多行为协调;对不同的动作价值函数进行简单的求和,根据情况切换动作价值函数,用之前得到的动作价值函数作为新动作价值函数的初始值进行学习。将球射进球门,避免与敌人发生碰撞。该任务可分解为射门子任务和避碰子任务。这些子任务应该同时完成,但不是相互独立的。<<ETX>>
A method is proposed which accomplishes a whole task consisting of plural subtasks by coordinating multiple behaviors acquired by a vision-based reinforcement learning. First, individual behaviors which achieve the corresponding subtasks are independently acquired by Q-learning, a widely used reinforcement learning method. Each learned behavior can be represented by an action-value function in terms of state of the environment and robot action. Next, three kinds of coordinations of multiple behaviors are considered; simple summation of different action-value functions, switching action-value functions according to situations, and learning with previously obtained action-value functions as initial values of a new action-value function. A task of shooting a ball into the goal avoiding collisions with an enemy is examined. The task can be decomposed into a ball shooting subtask and a collision avoiding subtask. These subtasks should be accomplished simultaneously, but they are not independent of each other.<<ETX>>