Learning Routines for Effective Off-Policy Reinforcement Learning

Learning Routines for Effective Off-Policy Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Edoardo Cetin;O. Çeliktutan
Edoardo Cetin;O. Çeliktutan
中科院分区:
其他
文献类型:
--
作者:
Edoardo Cetin;O. Çeliktutan

文献摘要

相似文献

强化学习的性能取决于设计适当的动作空间,其中每个动作的效果都是可测量的,但足够细化以允许灵活的行为。到目前为止,这个过程涉及到用户在可用操作及其执行频率方面的重要选择。我们提出了一种新颖的强化学习框架,可以有效地解除这些限制。在我们的框架内,代理在例程空间中学习有效的行为:一个新的、更高级别的动作空间,其中每个例程代表一组任意长度的“等效”粒状动作序列。我们的日常空间是端到端学习的,以促进实现潜在的非策略强化学习目标。我们将我们的框架应用于两种最先进的离策略算法,并表明最终的代理获得了相关的性能改进,同时每集与环境的交互更少,从而提高了计算效率。
The performance of reinforcement learning depends upon designing an appropriate action space, where the effect of each action is measurable, yet, granular enough to permit flexible behavior. So far, this process involved non-trivial user choices in terms of the available actions and their execution frequency. We propose a novel framework for reinforcement learning that effectively lifts such constraints. Within our framework, agents learn effective behavior over a routine space: a new, higher-level action space, where each routine represents a set of 'equivalent' sequences of granular actions with arbitrary length. Our routine space is learned end-to-end to facilitate the accomplishment of underlying off-policy reinforcement learning objectives. We apply our framework to two state-of-the-art off-policy algorithms and show that the resulting agents obtain relevant performance improvements while requiring fewer interactions with the environment per episode, improving computational efficiency.