Adversarial Intrinsic Motivation for Reinforcement Learning

Adversarial Intrinsic Motivation for Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Ishan Durugkar;Mauricio Tec;S. Niekum;P. Stone
Ishan Durugkar;Mauricio Tec;S. Niekum;P. Stone
中科院分区:
其他
文献类型:
--
作者:
Ishan Durugkar;Mauricio Tec;S. Niekum;P. Stone

文献摘要

被引文献

相似文献

以最大程度地减少参考分布的不匹配的学习,已证明对生成建模和模仿学习有用。在本文中,我们调查了这样一个目标,即政策的国家访问分配与目标分配之间的Wasserstein-1距离是否可以有效地用于加强学习(RL)任务。具体而言,本文着重于目标条件的强化学习,其中理想化的(无法实现的)目标分布在目标方面具有完全的衡量标准。本文介绍了Markov决策过程(MDP)特异性的准学,并使用此准计来估计上述Wasserstein-1距离。它进一步表明,最小化此Wasserstein-1距离的政策是在尽可能少的步骤中实现目标的政策。我们的方法称为对抗性的内在动机(AIM),估计了Wasserstein-1通过其双重目标估算的距离,并使用它来计算补充奖励功能。我们的实验表明,此奖励函数在MDP中的过渡方面顺利改变,并指导代理商的探索以有效地找到目标。此外,与其他鼓励探索或加速学习的奖励相比,我们将目标与事后的经验重播(她)相结合,并表明所产生的算法在几个模拟的机器人技术任务上大大加速了学习。
Learning with an objective to minimize the mismatch with a reference distribution has been shown to be useful for generative modeling and imitation learning. In this paper, we investigate whether one such objective, the Wasserstein-1 distance between a policy's state visitation distribution and a target distribution, can be utilized effectively for reinforcement learning (RL) tasks. Specifically, this paper focuses on goal-conditioned reinforcement learning where the idealized (unachievable) target distribution has full measure at the goal. This paper introduces a quasimetric specific to Markov Decision Processes (MDPs) and uses this quasimetric to estimate the above Wasserstein-1 distance. It further shows that the policy that minimizes this Wasserstein-1 distance is the policy that reaches the goal in as few steps as possible. Our approach, termed Adversarial Intrinsic Motivation (AIM), estimates this Wasserstein-1 distance through its dual objective and uses it to compute a supplemental reward function. Our experiments show that this reward function changes smoothly with respect to transitions in the MDP and directs the agent's exploration to find the goal efficiently. Additionally, we combine AIM with Hindsight Experience Replay (HER) and show that the resulting algorithm accelerates learning significantly on several simulated robotics tasks when compared to other rewards that encourage exploration or accelerate learning.