Learning with Options that Terminate Off-Policy

Learning with Options that Terminate Off-Policy
复制标题

使用终止偏离策略的选项进行学习

DOI:
10.1609/aaai.v32i1.11740
复制
发表时间:
2017
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Nowé
A. Nowé
中科院分区:
--
文献类型:
--
作者:
A. Harutyunyan;Peter Vrancx;Pierre;Doina Precup;A. Nowé

文献摘要

被引文献

相似文献

时间抽象的行为或期权由策略和终止条件指定:策略指导期权行为,终止条件大致确定其长度。一般来说,用更长的选项学习(比如用多步返回学习)被认为更有效率。然而,如果为任务设置的选项不是理想的,并且不能很好地表达原始的最优策略,则较短的选项提供了更多的灵活性,并且可以产生更好的解。因此,终止条件使学习效率与解的质量不一致。我们建议通过将行为和目标终止分离来解决这一困境,就像在非策略学习中使用策略一样。为此,我们给出了一个新的算法Q(Beta),它学习关于任何终止条件的解,而不管期权实际上是如何终止的。我们通过将带选项的学习投射到一个公共框架中来获得Q(Beta),该框架具有经过充分研究的多步策略学习。我们对我们的算法进行了实证验证,并证明了它是正确的。
A temporally abstract action, or an option, is specified by a policy and a termination condition: the policy guides the option behavior, and the termination condition roughly determines its length. Generally, learning with longer options (like learning with multi-step returns) is known to be more efficient. However, if the option set for the task is not ideal, and cannot express the primitive optimal policy well, shorter options offer more flexibility and can yield a better solution. Thus, the termination condition puts learning efficiency at odds with solution quality. We propose to resolve this dilemma by decoupling the behavior and target terminations, just like it is done with policies in off-policy learning. To this end, we give a new algorithm, Q(beta), that learns the solution with respect to any termination condition, regardless of how the options actually terminate. We derive Q(beta) by casting learning with options into a common framework with well-studied multi-step off policy learning. We validate our algorithm empirically, and show that it holds up to its motivating claims.