Off-Policy Deep Reinforcement Learning with Analogous Disentangled Exploration

Off-Policy Deep Reinforcement Learning with Analogous Disentangled Exploration
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
--
影响因子:
--
通讯作者:
Anji Liu;Yitao Liang;Guy Van den Broeck
Anji Liu;Yitao Liang;Guy Van den Broeck
中科院分区:
其他
文献类型:
--
作者:
Anji Liu;Yitao Liang;Guy Van den Broeck

文献摘要

相似文献

非策略强化学习(Off-policy reinforcement learning, RL)关注的是通过执行另一个收集经验样本的策略来学习一个奖励策略。虽然前一个策略(即目标策略)是有益的,但在表达中(在大多数情况下,是确定性的),但在后一个任务中表现良好,相反,需要一个提供指导和有效探索的表达策略(即行为策略)。与大多数在最优性和表达性之间进行权衡的方法相反,解纠缠框架显式地将两个目标解耦,每个目标都由不同的单独策略处理。尽管可以根据自己的目标自由地设计和优化这两个策略,但天真地将它们分开可能会导致低效的学习或稳定性问题。为了缓解这一问题,我们提出了类似解纠缠的演员-评论家(ADAC)方法,设计了类似的演员和评论家对。具体来说,ADAC利用Stein变分梯度下降(SVGD)的一个关键特性来约束基于表达能量的行为策略相对于目标策略进行有效的探索。此外,引入了一个类似的批评对,以原则性的方式纳入内在奖励,并在理论上保证了整体学习的稳定性和有效性。我们对14个连续控制任务的仅环境奖励的ADAC进行了实证评估,并报告了其中10个任务的最新进展。我们进一步证明,当ADAC与内在奖励配对时,在探索挑战性任务中优于替代方案。
Off-policy reinforcement learning (RL) is concerned with learning a rewarding policy by executing another policy that gathers samples of experience. While the former policy (i.e. target policy) is rewarding but in-expressive (in most cases, deterministic), doing well in the latter task, in contrast, requires an expressive policy (i.e. behavior policy) that offers guided and effective exploration. Contrary to most methods that make a trade-off between optimality and expressiveness, disentangled frameworks explicitly decouple the two objectives, which each is dealt with by a distinct separate policy. Although being able to freely design and optimize the two policies with respect to their own objectives, naively disentangling them can lead to inefficient learning or stability issues. To mitigate this problem, our proposed method Analogous Disentangled Actor-Critic (ADAC) designs analogous pairs of actors and critics. Specifically, ADAC leverages a key property about Stein variational gradient descent (SVGD) to constraint the expressive energy-based behavior policy with respect to the target one for effective exploration. Additionally, an analogous critic pair is introduced to incorporate intrinsic rewards in a principled manner, with theoretical guarantees on the overall learning stability and effectiveness. We empirically evaluate environment-reward-only ADAC on 14 continuous-control tasks and report the state-of-the-art on 10 of them. We further demonstrate ADAC, when paired with intrinsic rewards, outperform alternatives in exploration-challenging tasks.