Speeding-up Reinforcement Learning with Multi-step Actions

Speeding-up Reinforcement Learning with Multi-step Actions
复制标题

通过多步骤行动加速强化学习

DOI:
10.1007/3-540-46084-5_132
复制
发表时间:
2002
期刊:
ArXiv
影响因子:
--
通讯作者:
Martin A. Riedmiller
Martin A. Riedmiller
中科院分区:
--
文献类型:
--
作者:
Ralf Schoknecht;Martin A. Riedmiller

文献摘要

被引文献

相似文献

近年来,时间抽象的分层概念已被集成到强化学习框架中,以提高可扩展性。然而,现有的方法仅限于域分解成子任务是已知的先验。在本文中,我们提出的概念,明确选择时间尺度相关的行动,如果没有subgoalrelated抽象的行动。这是通过在不同时间尺度上的多步动作来实现的,这些动作组合在一个动作集中。在多SAQ学习算法中利用了动作集的特殊结构。通过同时在不同的明确指定的时间尺度上学习,可以实现学习速度的显著提高。这在两个基准问题上得到了证明。
In recent years hierarchical concepts of temporal abstraction have been integrated in the reinforcement learning framework to improve scalability. However, existing approaches are limited to domains where a decomposition into subtasks is known a priori. In this paper we propose the concept of explicitly selecting time scale related actions if no subgoalrelated abstract actions are available. This is realised with multistep actions on different time scales that are combined in one single action set. The special structure of the action set is exploited in the MSAQ-learning algorithm. By learning on different explicitly specified time scales simultaneously, a considerable improvement of learning speed can be achieved. This is demonstrated on two benchmark problems.