A Theory of Abstraction in Reinforcement Learning

A Theory of Abstraction in Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2203.00397
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
David Abel
David Abel
中科院分区:
其他
文献类型:
--
作者:
David Abel

文献摘要

被引文献

相似文献

强化学习定义了代理所面临的问题,这些代理仅通过行动和观察来学习做出正确的决策。为了成为有效的问题解决者,这些智能体必须有效地探索广阔的世界,从延迟反馈中分配信用,并推广到新的经验,同时利用有限的数据,计算资源和感知带宽。抽象对于所有这些努力都是必不可少的。通过抽象,智能体可以形成其环境的简洁模型,这些模型支持理性的、自适应的决策者所需的许多实践。在这篇论文中,我提出了强化学习中的抽象理论。我首先为执行抽象过程的函数提供三个必要条件:它们应该1)保持接近最优行为的表示,2)有效地学习和构建,3)减少规划或学习时间。然后,我提出了一套新的算法和分析,阐明了代理如何根据这些必要条件学习抽象。总的来说,这些结果为发现和使用抽象提供了一条部分路径,从而最大限度地降低了有效强化学习的复杂性。
Reinforcement learning defines the problem facing agents that learn to make good decisions through action and observation alone. To be effective problem solvers, such agents must efficiently explore vast worlds, assign credit from delayed feedback, and generalize to new experiences, all while making use of limited data, computational resources, and perceptual bandwidth. Abstraction is essential to all of these endeavors. Through abstraction, agents can form concise models of their environment that support the many practices required of a rational, adaptive decision maker. In this dissertation, I present a theory of abstraction in reinforcement learning. I first offer three desiderata for functions that carry out the process of abstraction: they should 1) preserve representation of near-optimal behavior, 2) be learned and constructed efficiently, and 3) lower planning or learning time. I then present a suite of new algorithms and analysis that clarify how agents can learn to abstract according to these desiderata. Collectively, these results provide a partial path toward the discovery and use of abstraction that minimizes the complexity of effective reinforcement learning.