Efficient exploration with Double Uncertain Value Networks

Efficient exploration with Double Uncertain Value Networks
复制标题

利用双重不确定价值网络进行高效探索

DOI:
--
复制
发表时间:
2017
期刊:
arXiv.org
影响因子:
--
通讯作者:
C. Jonker
C. Jonker
中科院分区:
--
文献类型:
--
作者:
T. Moerland;J. Broekens;C. Jonker

文献摘要

被引文献

相似文献

本文通过跟踪每个可用动作的值的不确定性来研究强化学习代理的定向探索。我们确定了两个与勘探相关的不确定性来源。前者源于有限的数据(参数不确定性),后者源于收益率的分布(收益率不确定性)。我们确定了用深度神经网络学习这些分布的方法,其中我们用贝叶斯丢弃来估计参数不确定性,而回报不确定性则以高斯分布通过Bellman方程传播。然后,我们证明了这两者可以在一个网络中联合估计,我们称之为双重不确定价值网络。该策略直接从基于Thompson抽样的学习分布中得到。实验结果表明,这两种类型的不确定性都可以极大地改善具有强大探索挑战的领域的学习。
This paper studies directed exploration for reinforcement learning agents by tracking uncertainty about the value of each available action. We identify two sources of uncertainty that are relevant for exploration. The first originates from limited data (parametric uncertainty), while the second originates from the distribution of the returns (return uncertainty). We identify methods to learn these distributions with deep neural networks, where we estimate parametric uncertainty with Bayesian drop-out, while return uncertainty is propagated through the Bellman equation as a Gaussian distribution. Then, we identify that both can be jointly estimated in one network, which we call the Double Uncertain Value Network. The policy is directly derived from the learned distributions based on Thompson sampling. Experimental results show that both types of uncertainty may vastly improve learning in domains with a strong exploration challenge.