MADE: Exploration via Maximizing Deviation from Explored Regions

MADE: Exploration via Maximizing Deviation from Explored Regions
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Tianjun Zhang;Paria Rashidinejad;Jiantao Jiao;Yuandong Tian;Joseph Gonzalez;Stuart J. Russell
Tianjun Zhang;Paria Rashidinejad;Jiantao Jiao;Yuandong Tian;Joseph Gonzalez;Stuart J. Russell
中科院分区:
其他
文献类型:
--
作者:
Tianjun Zhang;Paria Rashidinejad;Jiantao Jiao;Yuandong Tian;Joseph Gonzalez;Stuart J. Russell

文献摘要

相似文献

在在线强化学习(RL)中,有效的探索在具有稀疏奖励的高维环境中仍然特别具有挑战性。在低维环境中,表格参数化是可能的,基于计数的置信上限(UCB)探索方法实现极小极大接近最优率。然而,目前还不清楚如何有效地实现UCB在现实的RL任务,涉及非线性函数逼近。为了解决这个问题,我们提出了一种新的探索方法,通过\textit{最大化}的偏差,从探索区域的下一个政策的占用。我们将此项作为自适应正则化器添加到标准RL目标中,以平衡探索与利用。我们将新的目标与可证明收敛的算法配对,从而产生一个新的内在奖励,以调整现有的奖金。所提出的内在奖励易于实现,并与其他现有的强化学习算法进行联合收割机进行探索。作为概念证明,我们评估了各种基于模型和无模型算法的表格示例的新内在奖励,显示了对仅计数探索策略的改进。当对MiniGrid和DeepMind Control Suite基准测试的导航和运动任务进行测试时,我们的方法比最先进的方法显着提高了样本效率。我们的代码可在https://github.com/tianjunz/MADE上获得。
In online reinforcement learning (RL), efficient exploration remains particularly challenging in high-dimensional environments with sparse rewards. In low-dimensional environments, where tabular parameterization is possible, count-based upper confidence bound (UCB) exploration methods achieve minimax near-optimal rates. However, it remains unclear how to efficiently implement UCB in realistic RL tasks that involve non-linear function approximation. To address this, we propose a new exploration approach via \textit{maximizing} the deviation of the occupancy of the next policy from the explored regions. We add this term as an adaptive regularizer to the standard RL objective to balance exploration vs. exploitation. We pair the new objective with a provably convergent algorithm, giving rise to a new intrinsic reward that adjusts existing bonuses. The proposed intrinsic reward is easy to implement and combine with other existing RL algorithms to conduct exploration. As a proof of concept, we evaluate the new intrinsic reward on tabular examples across a variety of model-based and model-free algorithms, showing improvements over count-only exploration strategies. When tested on navigation and locomotion tasks from MiniGrid and DeepMind Control Suite benchmarks, our approach significantly improves sample efficiency over state-of-the-art methods. Our code is available at https://github.com/tianjunz/MADE.