First-Order Function Approximation for Transfer Learning in Relational MDPs

First-Order Function Approximation for Transfer Learning in Relational MDPs
复制标题

关系 MDP 中迁移学习的一阶函数逼近

DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Ronald P. A. Petrick
Ronald P. A. Petrick
中科院分区:
--
文献类型:
--
作者:
Jun Hao Alvin Ng;Ronald P. A. Petrick

文献摘要

被引文献

相似文献

具有一阶结构的规划问题可以用关系马尔可夫决策过程(rmdp)紧凑地建模。如果模型是未知的,可以使用基于值的强化学习方法来解决这些问题。动作值函数用合态基态流的特征来近似。然而,这种近似并没有利用RMDP的一阶结构,生成的策略只能求解RMDP的一个基MDP。我们的目标是学习一种广义函数近似,它可以导出一个可以解决多个地面mdp的策略。我们通过使用连接提升状态流作为一阶特征来实现这一点。这种一阶近似提供了更好的泛化,但粒度较粗,可能会降低性能。我们提出将一阶地物与地面地物相结合,以发挥两者的优势。四个领域的经验结果表明,我们的方法可以推广到问题上,无论其规模如何,并允许迁移学习。
Planning problems with a first-order structure can be modelled compactly with Relational Markov Decision Processes (RMDPs). If the model is unknown, value-based reinforcement learning methods can be used to solve these problems. The action-value function is approximated with features which are conjunctive ground state fluents. However, this approximation does not exploit the first-order structure of RMDPs and the generated policy can only solve a ground MDP of the RMDP. Our objective is to learn a generalised function approximation which induces a policy that can solve multiple ground MDPs. We achieve this by using conjunctive lifted state fluents as first-order features. This first-order approximation gives better generalisation but has a coarser granularity which can worsen performance. We propose the combination of first-order features and ground features to get both of their strengths. Empirical results for four domains show that our method could generalise over problems regardless of their scales and allow transfer learning.