First-Order Function Approximation for Transfer Learning in Relational MDPs
First-Order Function Approximation for Transfer Learning in Relational MDPs
复制标题
关系 MDP 中迁移学习的一阶函数逼近
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Ronald P. A. Petrick
中科院分区:
文献类型:
--
作者:
Jun Hao Alvin Ng;Ronald P. A. Petrick
Planning problems with a first-order structure can be modelled compactly with Relational Markov Decision Processes (RMDPs). If the model is unknown, value-based reinforcement learning methods can be used to solve these problems. The action-value function is approximated with features which are conjunctive ground state fluents. However, this approximation does not exploit the first-order structure of RMDPs and the generated policy can only solve a ground MDP of the RMDP. Our objective is to learn a generalised function approximation which induces a policy that can solve multiple ground MDPs. We achieve this by using conjunctive lifted state fluents as first-order features. This first-order approximation gives better generalisation but has a coarser granularity which can worsen performance. We propose the combination of first-order features and ground features to get both of their strengths. Empirical results for four domains show that our method could generalise over problems regardless of their scales and allow transfer learning.