Symbolic Network: Generalized Neural Policies for Relational MDPs

Symbolic Network: Generalized Neural Policies for Relational MDPs
复制标题

符号网络:关系 MDP 的广义神经策略

DOI:
--
复制
发表时间:
2020
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
Mausam
Mausam
中科院分区:
--
文献类型:
--
作者:
Sankalp Garg;Aniket Bajpai;Mausam

文献摘要

被引文献

相似文献

关系马尔可夫决策过程(Relational Markov Decision Process,RMDP)是一种一阶表示方法,用于表示单个概率规划域的所有实例,其中对象的数量可能是无限的。RMDPs的早期工作输出广义(实例独立)一阶策略或值函数,作为一种手段来解决一个域的所有实例。不幸的是,由于在这些策略或值函数中使用的表示空间的固有限制,这一系列工作取得了有限的成功。神经模型是否可以通过轻松地表示更复杂的广义策略来提供缺失的环节,从而使它们对给定域的所有实例都有效? 我们提出了SymNet,第一个神经方法来解决RMDPs的概率规划语言的RMDPs。SymNet使用来自域的训练实例为Rounds域训练一组共享参数。对于每个实例,SymNet首先将其转换为实例图,然后使用关系神经模型来计算节点嵌入。然后,它将每个地面动作作为一阶动作符号和与动作相关的节点嵌入的函数进行评分。给定来自相同域的新测试实例,具有预训练参数的SymNet架构对每个地面动作进行评分并选择最佳动作。这可以在单个前向传递中完成,而无需对测试实例进行任何重新训练,从而隐式地表示整个域的神经广义策略。我们在IPPC的九个Rendezvous域上的实验表明,SymNet策略明显优于随机策略,有时甚至比从头开始训练最先进的深度反应策略更有效。
A Relational Markov Decision Process (RMDP) is a first-order representation to express all instances of a single probabilistic planning domain with possibly unbounded number of objects. Early work in RMDPs outputs generalized (instance-independent) first-order policies or value functions as a means to solve all instances of a domain at once. Unfortunately, this line of work met with limited success due to inherent limitations of the representation space used in such policies or value functions. Can neural models provide the missing link by easily representing more complex generalized policies, thus making them effective on all instances of a given domain? We present SymNet, the first neural approach for solving RMDPs that are expressed in the probabilistic planning language of RDDL. SymNet trains a set of shared parameters for an RDDL domain using training instances from that domain. For each instance, SymNet first converts it to an instance graph and then uses relational neural models to compute node embeddings. It then scores each ground action as a function over the first-order action symbols and node embeddings related to the action. Given a new test instance from the same domain, SymNet architecture with pre-trained parameters scores each ground action and chooses the best action. This can be accomplished in a single forward pass without any retraining on the test instance, thus implicitly representing a neural generalized policy for the whole domain. Our experiments on nine RDDL domains from IPPC demonstrate that SymNet policies are significantly better than random and sometimes even more effective than training a state-of-the-art deep reactive policy from scratch.