Two-Timescale Networks for Nonlinear Value Function Approximation

Two-Timescale Networks for Nonlinear Value Function Approximation
复制标题

用于非线性值函数逼近的双时间尺度网络

DOI:
10.7939/r3-dx5r-7020
复制
发表时间:
2019
期刊:
ArXiv
影响因子:
--
通讯作者:
Martha White
Martha White
中科院分区:
--
文献类型:
--
作者:
Wesley Chung;Somjit Nath;A. Joseph;Martha White

文献摘要

被引文献

相似文献

许多强化学习代理的关键组件是学习值函数,用于策略评估或控制。然而,许多学习值的算法都是为线性函数近似而设计的-具有固定的基或固定的表示。虽然非线性函数近似有一些合理的扩展,如非线性梯度时间差学习,但这些方法在很大程度上没有被采用,而是避开了更简单但不合理的方法,如时间差学习和Q学习。在这项工作中,我们提供了一个双时标网络(TTN)架构,使线性方法可以用来学习值,而非线性表示在较慢的时标学习。该方法有利于使用的线性设置,如数据效率最小二乘法,资格跟踪和最近开发的线性政策评估算法的无数的算法,提供非线性价值估计。我们证明了TTN的收敛性,特别注意确保快速线性分量在学习表示提供的潜在依赖特征下的收敛性。我们经验证明TTN的好处,相比其他非线性值函数逼近算法,无论是政策评估和控制。
A key component for many reinforcement learning agents is to learn a value function, either for policy evaluation or control. Many of the algorithms for learning values, however, are designed for linear function approximation—with a fixed basis or fixed representation. Though there have been a few sound extensions to nonlinear function approximation, such as nonlinear gradient temporal difference learning, these methods have largely not been adopted, eschewed in favour of simpler but not sound methods like temporal difference learning and Q-learning. In this work, we provide a two-timescale network (TTN) architecture that enables linear methods to be used to learn values, with a nonlinear representation learned at a slower timescale. The approach facilitates the use of algorithms developed for the linear setting, such as data-efficient least-squares methods, eligibility traces and the myriad of recently developed linear policy evaluation algorithms, to provide nonlinear value estimates. We prove convergence for TTNs, with particular care given to ensure convergence of the fast linear component under potentially dependent features provided by the learned representation. We empirically demonstrate the benefits of TTNs, compared to other nonlinear value function approximation algorithms, both for policy evaluation and control.