Exploration in relational domains for model-based reinforcement learning

Exploration in relational domains for model-based reinforcement learning
复制标题

基于模型的强化学习的关系领域探索

DOI:
10.5555/2503308.2503360
复制
发表时间:
2012
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
K. Kersting
K. Kersting
中科院分区:
--
文献类型:
--
作者:
Tobias Lang;Marc Toussaint;K. Kersting

文献摘要

被引文献

相似文献

强化学习的一个基本问题是平衡探索和剥削。我们通过开发E3和R-MAX算法概念的关系扩展,在大型随机关系域中基于模型的增强学习的背景下解决这个问题。在指数级的状态空间中有效探索需要利用学习模型的概括:在命题环境中,什么将被视为一种新颖的情况,值得探索的关系可能是在关系环境中是一个众所周知的环境,在这种情况下,剥削是有希望的。为了解决这个问题,我们介绍了关系计数函数,该功能概括了国家和行动访问计数的经典概念。在我们拥有关系kwik学习者和近乎最佳计划者的假设下,我们使用计数函数对框架的探索效率提供了保证。我们提出了一种具体的探索算法,该算法将实践上有效的概率规则学习者和关系计划者(不可保证)整合在一起,并采用了学到的关系规则的环境,以建模国家和行动的新颖性。我们在嘈杂的3D模拟机器人操纵问题和国际规划竞赛领域的结果表明,我们的方法比现有的命题更有效,并且是有效的探索技术。
A fundamental problem in reinforcement learning is balancing exploration and exploitation. We address this problem in the context of model-based reinforcement learning in large stochastic relational domains by developing relational extensions of the concepts of the E3 and R-MAX algorithms. Efficient exploration in exponentially large state spaces needs to exploit the generalization of the learned model: what in a propositional setting would be considered a novel situation and worth exploration may in the relational setting be a well-known context in which exploitation is promising. To address this we introduce relational count functions which generalize the classical notion of state and action visitation counts. We provide guarantees on the exploration efficiency of our framework using count functions under the assumption that we had a relational KWIK learner and a near-optimal planner. We propose a concrete exploration algorithm which integrates a practically efficient probabilistic rule learner and a relational planner (for which there are no guarantees, however) and employs the contexts of learned relational rules as features to model the novelty of states and actions. Our results in noisy 3D simulated robot manipulation problems and in domains of the international planning competition demonstrate that our approach is more effective than existing propositional and factored exploration techniques.