Action-Gap Phenomenon in Reinforcement Learning

Action-Gap Phenomenon in Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2011-12
期刊:
--
影响因子:
--
通讯作者:
A. Farahmand
A. Farahmand
中科院分区:
其他
文献类型:
--
作者:
A. Farahmand

文献摘要

被引文献

相似文献

许多强化学习问题的实践者已经观察到,即使估计的(动作-)值函数仍远未达到最优值,智能体的性能通常也会非常接近最优性能。本文的目的是通过引入作用间隙规律性的概念来解释和形式化这一现象。作为一个典型的结果,我们证明了对于一个遵循贪婪策略\(\hat{\pi}\)的智能体对于一个动作-值函数$\(\hat{Q}\)$,其性能损失$\(E[V^*(X) - V^{\hat{X}} (X)]\)$的上界为$\(O(|| \hat{Q} - Q^*||_\infty^{1+\zeta}\))$,其中ζ≥= 0)是量化动作-间隙规律性的参数。对于ζ > 0,我们的结果表明,与之前的分析结果相比,性能损失较小。最后,我们展示了这种规律性如何影响近似迭代算法族的性能。
Many practitioners of reinforcement learning problems have observed that oftentimes the performance of the agent reaches very close to the optimal performance even though the estimated (action-)value function is still far from the optimal one. The goal of this paper is to explain and formalize this phenomenon by introducing the concept of the action-gap regularity. As a typical result, we prove that for an agent following the greedy policy \(\hat{\pi}\) with respect to an action-value function $\(\hat{Q}\)$, the performance loss $\(E[V^*(X) - V^{\hat{X}} (X)]\)$ is upper bounded by $\(O(|| \hat{Q} - Q^*||_\infty^{1+\zeta}\))$, in which ζ ≥ = 0) is the parameter quantifying the action-gap regularity. For ζ > 0, our results indicate smaller performance loss compared to what previous analyses had suggested. Finally, we show how this regularity affects the performance of the family of approximate value iteration algorithms.