Instance based Generalization in Reinforcement Learning

Instance based Generalization in Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Martín Bertrán;Natalia Martínez;Mariano Phielipp;G. Sapiro
Martín Bertrán;Natalia Martínez;Mariano Phielipp;G. Sapiro
中科院分区:
其他
文献类型:
--
作者:
Martín Bertrán;Natalia Martínez;Mariano Phielipp;G. Sapiro

文献摘要

相似文献

通过深度强化学习(RL)训练的智能体通常无法泛化到看不见的环境,即使这些环境与训练水平具有相同的潜在动态。理解RL的泛化特性是现代机器学习的挑战之一。为了实现这一目标,我们在部分可观察马尔可夫决策过程(POMDPs)的背景下分析策略学习,并将训练水平的动态形式化为实例。我们证明,独立的探索策略,重用实例引入了显着的变化,有效的马尔可夫动态代理在训练过程中观察。最大化预期奖励的影响,通过诱导不希望的实例特定的speedrunning政策,而不是一般化的,这是次优的训练集上的代理学习的信念状态。我们根据训练实例的数量为训练和测试环境中的值差距提供泛化范围,并使用基于这些的见解来提高看不见的水平上的性能。我们建议训练一个共享的信念表示在一个合奏的专门政策,从中我们计算一个共识的政策,用于数据收集,不允许实例特定的剥削。我们通过实验验证了我们的理论,观察结果以及CoinRun基准测试中提出的计算解决方案。
Agents trained via deep reinforcement learning (RL) routinely fail to generalize to unseen environments, even when these share the same underlying dynamics as the training levels. Understanding the generalization properties of RL is one of the challenges of modern machine learning. Towards this goal, we analyze policy learning in the context of Partially Observable Markov Decision Processes (POMDPs) and formalize the dynamics of training levels as instances. We prove that, independently of the exploration strategy, reusing instances introduces significant changes on the effective Markov dynamics the agent observes during training. Maximizing expected rewards impacts the learned belief state of the agent by inducing undesired instance specific speedrunning policies instead of generalizeable ones, which are suboptimal on the training set. We provide generalization bounds to the value gap in train and test environments based on the number of training instances, and use insights based on these to improve performance on unseen levels. We propose training a shared belief representation over an ensemble of specialized policies, from which we compute a consensus policy that is used for data collection, disallowing instance specific exploitation. We experimentally validate our theory, observations, and the proposed computational solution over the CoinRun benchmark.