Model Selection in Reinforcement Learning with General Function Approximations

Model Selection in Reinforcement Learning with General Function Approximations
复制标题

具有一般函数逼近的强化学习中的模型选择

DOI:
10.48550/arxiv.2207.02992
复制
发表时间:
2022
期刊:
Teratology
影响因子:
--
通讯作者:
Sayak Ray Chowdhury
Sayak Ray Chowdhury
中科院分区:
--
文献类型:
--
作者:
Avishek Ghosh;Sayak Ray Chowdhury

文献摘要

参考文献

被引文献

相似文献

我们考虑在一般函数近似下经典强化学习 (RL) 环境的模型选择——多武装老虎机 (MAB) 和马尔可夫决策过程 (MDP)。在模型选择框架中,我们不知道由 $\mathcal{F}$ 和 $\mathcal{M}$ 表示的函数类,它们分别是真实模型(MAB 的奖励生成函数和 MDP 的转换内核)所在的位置。相反,我们给出了 $M$ 嵌套函数(假设)类,以便真实模型至少包含在一个这样的类中。在本文中,我们提出并分析了 MAB 和 MDP 的有效模型选择算法,该算法\emph{适应}包含真实底层模型的最小函数类(在嵌套的 $M$ 类中)。在嵌套假设类的可分离性假设下,我们表明自适应算法的累积遗憾与先验知道正确函数类(即 $\cF$ 和 $\cM$)的预言机的累积遗憾相匹配。此外,对于这两种设置,我们表明模型选择的成本是后悔中的一个附加项,对学习范围 $T$ 具有弱(对数)依赖性。
We consider model selection for classic Reinforcement Learning (RL) environments -- Multi Armed Bandits (MABs) and Markov Decision Processes (MDPs) -- under general function approximations. In the model selection framework, we do not know the function classes, denoted by $\mathcal{F}$ and $\mathcal{M}$, where the true models -- reward generating function for MABs and and transition kernel for MDPs -- lie, respectively. Instead, we are given $M$ nested function (hypothesis) classes such that true models are contained in at-least one such class. In this paper, we propose and analyze efficient model selection algorithms for MABs and MDPs, that \emph{adapt} to the smallest function class (among the nested $M$ classes) containing the true underlying model. Under a separability assumption on the nested hypothesis classes, we show that the cumulative regret of our adaptive algorithms match to that of an oracle which knows the correct function classes (i.e., $\cF$ and $\cM$) a priori. Furthermore, for both the settings, we show that the cost of model selection is an additive term in the regret having weak (logarithmic) dependence on the learning horizon $T$.
DOI: --
发表时间: 2021
期刊: Proceedings of Machine Learning Research
影响因子: --
作者:
Arora, Raman;Vanislavov, Teodor Marinov;Mohri, Mehryar
通讯作者: Mohri, Mehryar