Time Limits in Reinforcement Learning

Time Limits in Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2017-12
期刊:
影响因子:
3.1
通讯作者:
Fabio Pardo;Arash Tavakoli;Vitaly Levdik;Petar Kormushev
Fabio Pardo;Arash Tavakoli;Vitaly Levdik;Petar Kormushev
中科院分区:
化学3区
文献类型:
--
作者:
Fabio Pardo;Arash Tavakoli;Vitaly Levdik;Petar Kormushev

文献摘要

被引文献

相似文献

在强化学习中,通常让代理与其环境交互固定的时间,然后重置它并在一系列情节中重复该过程。代理必须学习的任务可以是在(i)固定时期或(ii)无限期期间最大化其性能,其中时间限制仅在训练期间使用以使经验多样化。在本文中,我们正式说明了如何在这两种情况下有效处理时间限制,并解释了为什么不这样做会导致状态混叠和经验重放无效,从而导致策略不理想和训练不稳定。在情况(i)中,我们认为由于时间限制而导致的终止实际上是环境的一部分,因此剩余时间的概念应该作为代理输入的一部分包含在内,以避免违反马尔可夫性质。在情况(ii)中,时间限制不是环境的一部分,仅用于促进学习。我们认为,这种见解应该通过从每个部分事件结束时的状态价值引导来纳入。对于这两种情况,我们凭经验说明了我们考虑因素在提高现有强化学习算法的性能和稳定性方面的重要性,并在几个控制任务上展示了最先进的结果。
In reinforcement learning, it is common to let an agent interact for a fixed amount of time with its environment before resetting it and repeating the process in a series of episodes. The task that the agent has to learn can either be to maximize its performance over (i) that fixed period, or (ii) an indefinite period where time limits are only used during training to diversify experience. In this paper, we provide a formal account for how time limits could effectively be handled in each of the two cases and explain why not doing so can cause state-aliasing and invalidation of experience replay, leading to suboptimal policies and training instability. In case (i), we argue that the terminations due to time limits are in fact part of the environment, and thus a notion of the remaining time should be included as part of the agent's input to avoid violation of the Markov property. In case (ii), the time limits are not part of the environment and are only used to facilitate learning. We argue that this insight should be incorporated by bootstrapping from the value of the state at the end of each partial episode. For both cases, we illustrate empirically the significance of our considerations in improving the performance and stability of existing reinforcement learning algorithms, showing state-of-the-art results on several control tasks.