Online abstraction with MDP homomorphisms for Deep Learning

Online abstraction with MDP homomorphisms for Deep Learning
复制标题

DOI:
--
复制
发表时间:
2018-11
期刊:
--
影响因子:
--
通讯作者:
Ondrej Biza;Robert W. Platt
Ondrej Biza;Robert W. Platt
中科院分区:
其他
文献类型:
--
作者:
Ondrej Biza;Robert W. Platt

文献摘要

被引文献

相似文献

马尔可夫决策过程的抽象是解决复杂问题的有用工具,因为它可以忽略环境中不重要的方面,简化学习最优策略的过程。在本文中,我们提出了一种在具有连续状态空间的环境中查找抽象 MDP 的新算法。它基于 MDP 同态,即 MDP 之间的结构保留映射。我们展示了我们的算法从收集的经验中学习抽象的能力,并展示了如何重用抽象来指导代理遇到的新任务的探索。在我们的大多数实验中,我们新颖的任务转移方法优于基于深度 Q 网络的基线。源代码位于此 https URL。
Abstraction of Markov Decision Processes is a useful tool for solving complex problems, as it can ignore unimportant aspects of an environment, simplifying the process of learning an optimal policy. In this paper, we propose a new algorithm for finding abstract MDPs in environments with continuous state spaces. It is based on MDP homomorphisms, a structure-preserving mapping between MDPs. We demonstrate our algorithm's ability to learn abstractions from collected experience and show how to reuse the abstractions to guide exploration in new tasks the agent encounters. Our novel task transfer method outperforms baselines based on a deep Q-network in the majority of our experiments. The source code is at this https URL.