The Laplacian in RL: Learning Representations with Efficient Approximations

The Laplacian in RL: Learning Representations with Efficient Approximations
复制标题

强化学习中的拉普拉斯算子:通过高效近似学习表示

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Learning Representations
影响因子:
--
通讯作者:
Ofir Nachum
Ofir Nachum
中科院分区:
--
文献类型:
--
作者:
Yifan Wu;G. Tucker;Ofir Nachum

文献摘要

被引文献

相似文献

众所周知,图的拉普拉斯矩阵的最小特征向量提供了加权图的几何的简洁表示。在强化学习(RL)中,加权图可以被解释为由作用于环境的行为策略引起的状态转移过程,近似拉普拉斯算子的特征向量为状态表示学习提供了一种很有前途的方法。然而,现有的用于执行这种近似的方法不适合于一般的RL设置,主要原因有两个:第一,它们的计算代价很高,通常需要对大矩阵进行运算。其次,除了简单、表格、有限状态的设置之外,这些方法缺乏充分的理由。在这篇文章中,我们提出了一种完全通用和可伸缩的方法,在无模型的RL上下文中逼近拉普拉斯的特征向量。我们系统地评估了我们的方法,并经验表明,它超越了表格,有限状态的设置。即使在表格、有限状态设置中,其逼近特征向量的能力也优于以前的提议。最后,我们展示了在目标实现的RL任务中使用使用我们的方法学习的拉普拉斯表示的潜在好处,提供了我们的技术可以用于显著提高RL代理的性能的证据。
The smallest eigenvectors of the graph Laplacian are well-known to provide a succinct representation of the geometry of a weighted graph. In reinforcement learning (RL), where the weighted graph may be interpreted as the state transition process induced by a behavior policy acting on the environment, approximating the eigenvectors of the Laplacian provides a promising approach to state representation learning. However, existing methods for performing this approximation are ill-suited in general RL settings for two main reasons: First, they are computationally expensive, often requiring operations on large matrices. Second, these methods lack adequate justification beyond simple, tabular, finite-state settings. In this paper, we present a fully general and scalable method for approximating the eigenvectors of the Laplacian in a model-free RL context. We systematically evaluate our approach and empirically show that it generalizes beyond the tabular, finite-state setting. Even in tabular, finite-state settings, its ability to approximate the eigenvectors outperforms previous proposals. Finally, we show the potential benefits of using a Laplacian representation learned using our method in goal-achieving RL tasks, providing evidence that our technique can be used to significantly improve the performance of an RL agent.