Model Selection in Reinforcement Learning with General Function Approximations
Model Selection in Reinforcement Learning with General Function Approximations
复制标题
具有一般函数逼近的强化学习中的模型选择
DOI:
10.48550/arxiv.2207.02992
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Sayak Ray Chowdhury
中科院分区:
文献类型:
--
作者:
Avishek Ghosh;Sayak Ray Chowdhury
We consider model selection for classic Reinforcement Learning (RL) environments -- Multi Armed Bandits (MABs) and Markov Decision Processes (MDPs) -- under general function approximations. In the model selection framework, we do not know the function classes, denoted by $\mathcal{F}$ and $\mathcal{M}$, where the true models -- reward generating function for MABs and and transition kernel for MDPs -- lie, respectively. Instead, we are given $M$ nested function (hypothesis) classes such that true models are contained in at-least one such class. In this paper, we propose and analyze efficient model selection algorithms for MABs and MDPs, that \emph{adapt} to the smallest function class (among the nested $M$ classes) containing the true underlying model. Under a separability assumption on the nested hypothesis classes, we show that the cumulative regret of our adaptive algorithms match to that of an oracle which knows the correct function classes (i.e., $\cF$ and $\cM$) a priori. Furthermore, for both the settings, we show that the cost of model selection is an additive term in the regret having weak (logarithmic) dependence on the learning horizon $T$.
DOI:
--
发表时间:
2021
期刊:
Proceedings of Machine Learning Research
影响因子:
--
作者:
Arora, Raman;Vanislavov, Teodor Marinov;Mohri, Mehryar
通讯作者:
Mohri, Mehryar