Uncertainty-Aware Action Advising for Deep Reinforcement Learning Agents

Uncertainty-Aware Action Advising for Deep Reinforcement Learning Agents
复制标题

为深度强化学习代理提供不确定性感知行动建议

DOI:
--
复制
发表时间:
2020
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Matthew Taylor
Matthew Taylor
中科院分区:
--
文献类型:
--
作者:
Felipe Leno da Silva;Pablo Hernandez;Bilal Kartal;Matthew Taylor

文献摘要

被引文献

相似文献

虽然强化学习(RL)已经成为序列决策问题中最成功的学习方法之一,但RL技术的样本复杂性仍然是实际应用的主要挑战。为了应对这一挑战,每当一个主管政策(例如,遗留系统或人类演示者)可用时,代理可以利用来自该策略(建议)的样本来提高样本效率。然而,咨询意见通常是有限的,因此最好是针对代理人不确定应采取的最佳行动的国家。在这项工作中,我们提出了请求信心适度的政策建议(RCMP),一个行动建议框架,代理人要求咨询时,其认知的不确定性是高的某一状态。加拿大皇家骑警考虑到建议是有限的,可能是次优的。我们还描述了一种技术来估计代理的不确定性进行微小的修改,在标准值函数为基础的RL方法。我们的实证评估表明,RCMP执行更好的重要性建议,不接收建议,并在Gridworld和Atari乒乓场景中随机状态接收。
Although Reinforcement Learning (RL) has been one of the most successful approaches for learning in sequential decision making problems, the sample-complexity of RL techniques still represents a major challenge for practical applications. To combat this challenge, whenever a competent policy (e.g., either a legacy system or a human demonstrator) is available, the agent could leverage samples from this policy (advice) to improve sample-efficiency. However, advice is normally limited, hence it should ideally be directed to states where the agent is uncertain on the best action to execute. In this work, we propose Requesting Confidence-Moderated Policy advice (RCMP), an action-advising framework where the agent asks for advice when its epistemic uncertainty is high for a certain state. RCMP takes into account that the advice is limited and might be suboptimal. We also describe a technique to estimate the agent uncertainty by performing minor modifications in standard value-function-based RL methods. Our empirical evaluations show that RCMP performs better than Importance Advising, not receiving advice, and receiving it at random states in Gridworld and Atari Pong scenarios.