Finite-Time Performance Bounds and Adaptive Learning Rate Selection for Two Time-Scale Reinforcement Learning

Finite-Time Performance Bounds and Adaptive Learning Rate Selection for Two Time-Scale Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2019-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Harsh Gupta;R. Srikant;Lei Ying
Harsh Gupta;R. Srikant;Lei Ying
中科院分区:
其他
文献类型:
--
作者:
Harsh Gupta;R. Srikant;Lei Ying

文献摘要

被引文献

相似文献

我们研究了两种时间尺度线性随机逼近算法,它们可以用来模拟著名的强化学习算法,如GTD,GTD 2和TDC。我们提出了有限时间的学习率是固定的情况下的性能界限。在获得这些界限的关键思想是使用李雅普诺夫函数的奇异摄动理论的线性微分方程。我们使用的界限,设计一个自适应的学习率计划,显着提高了收敛速度在我们的实验中已知的最佳多项式衰减规则,并可用于潜在地提高性能的任何其他时间表,学习率在预定的时刻发生变化。
We study two time-scale linear stochastic approximation algorithms, which can be used to model well-known reinforcement learning algorithms such as GTD, GTD2, and TDC. We present finite-time performance bounds for the case where the learning rate is fixed. The key idea in obtaining these bounds is to use a Lyapunov function motivated by singular perturbation theory for linear differential equations. We use the bound to design an adaptive learning rate scheme which significantly improves the convergence rate over the known optimal polynomial decay rule in our experiments, and can be used to potentially improve the performance of any other schedule where the learning rate is changed at pre-determined time instants.