Optimal Rate of Convergence for Quasi-Stochastic Approximation.

Optimal Rate of Convergence for Quasi-Stochastic Approximation.
复制标题

准随机逼近的最佳收敛率。

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Sean P. Meyn
Sean P. Meyn
中科院分区:
--
文献类型:
--
作者:
A. Bernstein;Yue;Marcello Colombino;E. Dall’Anese;P. Mehta;Sean P. Meyn

文献摘要

参考文献

被引文献

相似文献

Robbins-Monro 随机逼近算法是强化学习 (RL) 的许多算法框架的基础,并且通常是解决(或近似解决)复杂最优控制问题的有效方法。然而,在许多情况下,由于固有的高方差,从业者无法应用这些技术。本文旨在为“准随机近似”提供一般基础,其中所考虑的所有过程都是确定性的,就像模拟中方差减少的准蒙特卡罗一样。根据算法中相关参数的调整,方差的减少可能会很大。本文引入了一种新的耦合参数,以在增益足够大的情况下建立最佳收敛速度。这些结果是针对线性模型建立的,并且还在非理想设置中进行了测试。这些一般结果的一个主要应用是用于确定性状态空间模型的一类新型强化学习算法。在这种情况下,主要贡献是一类算法,用于近似给定策略的价值函数,使用旨在引入探索的不同策略。
The Robbins-Monro stochastic approximation algorithm is a foundation of many algorithmic frameworks for reinforcement learning (RL), and often an efficient approach to solving (or approximating the solution to) complex optimal control problems. However, in many cases practitioners are unable to apply these techniques because of an inherent high variance. This paper aims to provide a general foundation for "quasi-stochastic approximation," in which all of the processes under consideration are deterministic, much like quasi-Monte-Carlo for variance reduction in simulation. The variance reduction can be substantial, subject to tuning of pertinent parameters in the algorithm. This paper introduces a new coupling argument to establish optimal rate of convergence provided the gain is sufficiently large. These results are established for linear models, and tested also in non-ideal settings. A major application of these general results is a new class of RL algorithms for deterministic state space models. In this setting, the main contribution is a class of algorithms for approximating the value function for a given policy, using a different policy designed to introduce exploration.
DOI: 10.1109/jiot.2018.2839563
发表时间: 2017-07
影响因子: 10.6
作者:
Tianyi Chen;G. Giannakis
通讯作者: Tianyi Chen;G. Giannakis