Optimistic Initialization for Exploration in Continuous Control

Optimistic Initialization for Exploration in Continuous Control
复制标题

DOI:
10.1609/aaai.v36i7.20727
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Sam Lobel;Omer Gottesman;Cameron S. Allen;Akhil Bagaria;G. Konidaris
Sam Lobel;Omer Gottesman;Cameron S. Allen;Akhil Bagaria;G. Konidaris
中科院分区:
其他
文献类型:
--
作者:
Sam Lobel;Omer Gottesman;Cameron S. Allen;Akhil Bagaria;G. Konidaris

文献摘要

相似文献

乐观初始化支持许多理论上合理的表格域探索方案;然而,在深层函数近似设置中,如果初始化天真,乐观可能很快消失。我们提出了一个框架,更有效地将乐观的初始化到强化学习的连续控制。我们的方法使用状态-动作空间的度量信息来估计哪些转换尚未探索,并明确保持相应状态-动作对的初始Q值乐观。我们还开发了有效地逼近这些训练目标的方法,并将领域知识纳入乐观的信封,以提高样本效率。我们经验性地评估这些方法在连续控制中的各种硬勘探问题,我们的方法优于现有的勘探技术。
Optimistic initialization underpins many theoretically sound exploration schemes in tabular domains; however, in the deep function approximation setting, optimism can quickly disappear if initialized naively. We propose a framework for more effectively incorporating optimistic initialization into reinforcement learning for continuous control. Our approach uses metric information about the state-action space to estimate which transitions are still unexplored, and explicitly maintains the initial Q-value optimism for the corresponding state-action pairs. We also develop methods for efficiently approximating these training objectives, and for incorporating domain knowledge into the optimistic envelope to improve sample efficiency. We empirically evaluate these approaches on a variety of hard exploration problems in continuous control, where our method outperforms existing exploration techniques.