Learning Value Functions from Undirected State-only Experience

Learning Value Functions from Undirected State-only Experience
复制标题

DOI:
10.48550/arxiv.2204.12458
复制
发表时间:
2022-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Matthew Chang;Arjun Gupta;Saurabh Gupta
Matthew Chang;Arjun Gupta;Saurabh Gupta
中科院分区:
其他
文献类型:
--
作者:
Matthew Chang;Arjun Gupta;Saurabh Gupta

文献摘要

相似文献

本文讨论了从无向纯状态经验(无动作标签的状态转移,即(S,S,r)元组)学习值函数的问题。我们首先从理论上刻画了Q-学习在这一背景下的适用性。我们证明了离散马尔可夫决策过程(MDP)中的表式Q-学习在任意的行动空间精化下学习相同的值函数。这一理论结果激励了潜在行动Q-学习或LAQ的设计,这是一种离线RL方法,可以从仅有状态的经验中学习有效的价值函数。潜在动作Q学习(LAQ)通过对通过潜在变量未来预测模型获得的离散潜在动作进行Q学习来学习值函数。我们证明了LAQ可以恢复与使用地面真实动作学习的值函数具有高度相关性的值函数。使用LAQ学习的值函数导致目标导向行为的样本有效获取,可以与特定于域的低级控制器一起使用,并且促进跨实施例的转移。我们在5个环境中的实验,从2D网格世界到真实环境中的3D视觉导航,展示了LAQ相对于更简单的替代方案、模仿学习先知和竞争方法的好处。
This paper tackles the problem of learning value functions from undirected state-only experience (state transitions without action labels i.e. (s,s',r) tuples). We first theoretically characterize the applicability of Q-learning in this setting. We show that tabular Q-learning in discrete Markov decision processes (MDPs) learns the same value function under any arbitrary refinement of the action space. This theoretical result motivates the design of Latent Action Q-learning or LAQ, an offline RL method that can learn effective value functions from state-only experience. Latent Action Q-learning (LAQ) learns value functions using Q-learning on discrete latent actions obtained through a latent-variable future prediction model. We show that LAQ can recover value functions that have high correlation with value functions learned using ground truth actions. Value functions learned using LAQ lead to sample efficient acquisition of goal-directed behavior, can be used with domain-specific low-level controllers, and facilitate transfer across embodiments. Our experiments in 5 environments ranging from 2D grid world to 3D visual navigation in realistic environments demonstrate the benefits of LAQ over simpler alternatives, imitation learning oracles, and competing methods.