Non-stationary Reinforcement Learning under General Function Approximation

Non-stationary Reinforcement Learning under General Function Approximation
复制标题

DOI:
10.48550/arxiv.2306.00861
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Songtao Feng;Ming Yin;Ruiquan Huang;Yu-Xiang Wang;J. Yang;Yitao Liang
Songtao Feng;Ming Yin;Ruiquan Huang;Yu-Xiang Wang;J. Yang;Yitao Liang
中科院分区:
其他
文献类型:
--
作者:
Songtao Feng;Ming Yin;Ruiquan Huang;Yu-Xiang Wang;J. Yang;Yitao Liang

文献摘要

相似文献

一般函数逼近是在广泛的强化学习(RL)场景中处理大的状态和动作空间的强大工具。然而,在一般函数近似下,对非平稳MDP的理论理解仍然有限。在本文中,我们进行了第一次这样的尝试。我们首先针对非平稳MDP提出了一种称为动态Bellman Eluder(DBE)维的新的复杂性度量,它包含了静态MDP和非平稳MDP中已有的大多数易处理的RL问题。基于所提出的复杂性度量,我们提出了一种新的基于置信度集的无模型算法SW-OPEA,该算法具有滑动窗口机制和一种新的针对非平稳MDP的置信度设计。然后,我们建立了该算法的动态后悔上界,并证明了只要变化预算不是很大,SW-OPEA就是可证明有效的。通过非平稳线性MDP和表格MDP的例子,我们进一步证明了我们的算法在小变化预算场景中的性能比现有的UCB类型的算法更好。据我们所知,这是第一次在非平稳MDP中用一般函数近似进行动态后悔分析。
General function approximation is a powerful tool to handle large state and action spaces in a broad range of reinforcement learning (RL) scenarios. However, theoretical understanding of non-stationary MDPs with general function approximation is still limited. In this paper, we make the first such an attempt. We first propose a new complexity metric called dynamic Bellman Eluder (DBE) dimension for non-stationary MDPs, which subsumes majority of existing tractable RL problems in static MDPs as well as non-stationary MDPs. Based on the proposed complexity metric, we propose a novel confidence-set based model-free algorithm called SW-OPEA, which features a sliding window mechanism and a new confidence set design for non-stationary MDPs. We then establish an upper bound on the dynamic regret for the proposed algorithm, and show that SW-OPEA is provably efficient as long as the variation budget is not significantly large. We further demonstrate via examples of non-stationary linear and tabular MDPs that our algorithm performs better in small variation budget scenario than the existing UCB-type algorithms. To the best of our knowledge, this is the first dynamic regret analysis in non-stationary MDPs with general function approximation.