Learning Macro-Actions in Reinforcement Learning

Learning Macro-Actions in Reinforcement Learning
复制标题

学习强化学习中的宏观动作

DOI:
--
复制
发表时间:
1998
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
J. Randløv
J. Randløv
中科院分区:
--
文献类型:
--
作者:
J. Randløv

文献摘要

被引文献

相似文献

我们提出了一种在强化学习过程中从原始动作自动构建宏观动作的方法。总的想法是,如果这样的行动模式得到了奖励,则强化在行动 a 之后执行行动 b 的倾向。我们在自行车任务、山上汽车任务、赛道任务和一些网格世界任务上测试了该方法。对于自行车和赛道任务,使用宏观动作大约使学习时间减少一半,而对于其中一项网格世界任务,学习时间减少了 5 倍。由于我们在结论中讨论的原因,该方法不适用于山上汽车任务。
We present a method for automatically constructing macro-actions from scratch from primitive actions during the reinforcement learning process. The overall idea is to reinforce the tendency to perform action b after action a if such a pattern of actions has been rewarded. We test the method on a bicycle task, the car-on-the-hill task, the race-track task and some grid-world tasks. For the bicycle and race-track tasks the use of macro-actions approximately halves the learning time, while for one of the grid-world tasks the learning time is reduced by a factor of 5. The method did not work for the car-on-the-hill task for reasons we discuss in the conclusion.