Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability

Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
Dibya Ghosh;Jad Rahme;Aviral Kumar;Amy Zhang;Ryan P. Adams;S. Levine
Dibya Ghosh;Jad Rahme;Aviral Kumar;Amy Zhang;Ryan P. Adams;S. Levine
中科院分区:
其他
文献类型:
--
作者:
Dibya Ghosh;Jad Rahme;Aviral Kumar;Amy Zhang;Ryan P. Adams;S. Levine

文献摘要

被引文献

相似文献

泛化是强化学习(RL)系统在真实的世界中部署的核心挑战。在本文中,我们证明了RL问题的顺序结构需要新的方法来推广,超越了监督学习中使用的研究充分的技术。虽然监督学习方法可以有效地推广,而无需明确考虑认知的不确定性,但我们表明,也许令人惊讶的是,在RL中并非如此。我们发现,从有限数量的训练条件推广到看不见的测试条件会导致隐式部分可观察性,甚至有效地将完全观察到的MDP转化为POMDP。根据这一观察,我们将强化学习中的泛化问题重新定义为解决诱导的部分可观察马尔可夫决策过程,我们称之为认知POMDP。我们证明了故障模式的算法,不适当地处理这部分的可观测性,并提出了一个简单的集成为基础的技术,近似解决部分观察到的问题。从经验上讲,我们证明了我们的简单算法来自认知POMDP实现了显着的收益,在推广目前的方法上的Procgen基准套件。
Generalization is a central challenge for the deployment of reinforcement learning (RL) systems in the real world. In this paper, we show that the sequential structure of the RL problem necessitates new approaches to generalization beyond the well-studied techniques used in supervised learning. While supervised learning methods can generalize effectively without explicitly accounting for epistemic uncertainty, we show that, perhaps surprisingly, this is not the case in RL. We show that generalization to unseen test conditions from a limited number of training conditions induces implicit partial observability, effectively turning even fully-observed MDPs into POMDPs. Informed by this observation, we recast the problem of generalization in RL as solving the induced partially observed Markov decision process, which we call the epistemic POMDP. We demonstrate the failure modes of algorithms that do not appropriately handle this partial observability, and suggest a simple ensemble-based technique for approximately solving the partially observed problem. Empirically, we demonstrate that our simple algorithm derived from the epistemic POMDP achieves significant gains in generalization over current methods on the Procgen benchmark suite.