Learning discrete state abstractions with deep variational inference

Learning discrete state abstractions with deep variational inference
复制标题

DOI:
--
复制
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Ondrej Biza;Robert W. Platt;Jan-Willem van de Meent;Lawson L. S. Wong
Ondrej Biza;Robert W. Platt;Jan-Willem van de Meent;Lawson L. S. Wong
中科院分区:
其他
文献类型:
--
作者:
Ondrej Biza;Robert W. Platt;Jan-Willem van de Meent;Lawson L. S. Wong

文献摘要

相似文献

在状态空间较大的领域中,抽象对于有效的顺序决策至关重要。在这项工作中,我们提出了一种学习近似互模拟的信息瓶颈方法,这是一种状态抽象。我们使用深度神经编码器将状态映射到连续嵌入。我们使用动作条件隐马尔可夫模型将这些嵌入映射到离散表示上,该模型与神经网络进行端到端的训练。我们的方法适用于高维状态的环境,并从马尔可夫决策过程中的代理收集的经验流中学习。通过这个已学习的离散抽象模型,我们可以在多目标强化学习环境中有效地计划未知的目标。我们在简化的带有图像状态的机器人操作领域中测试了我们的方法。我们还将其与以前的基于模型的方法进行了比较,以在离散的网格世界般的环境中找到互模拟。源代码可在此HTTPS URL上找到。
Abstraction is crucial for effective sequential decision making in domains with large state spaces. In this work, we propose an information bottleneck method for learning approximate bisimulations, a type of state abstraction. We use a deep neural encoder to map states onto continuous embeddings. We map these embeddings onto a discrete representation using an action-conditioned hidden Markov model, which is trained end-to-end with the neural network. Our method is suited for environments with high-dimensional states and learns from a stream of experience collected by an agent acting in a Markov decision process. Through this learned discrete abstract model, we can efficiently plan for unseen goals in a multi-goal Reinforcement Learning setting. We test our method in simplified robotic manipulation domains with image states. We also compare it against previous model-based approaches to finding bisimulations in discrete grid-world-like environments. Source code is available at this https URL.