Whittle index based Q-learning for restless bandits with average reward
Whittle index based Q-learning for restless bandits with average reward
复制标题
基于 Whittle 指数的 Q 学习,用于具有平均奖励的不安强盗
DOI:
10.1016/j.automatica.2022.110186
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
V. Borkar
中科院分区:
文献类型:
--
作者:
Konstantin Avrachenkov;V. Borkar
A novel reinforcement learning algorithm is introduced for multiarmed restless bandits with average reward, using the paradigms of Q-learning and Whittle index. Specifically, we leverage the structure of the Whittle index policy to reduce the search space of Q-learning, resulting in major computational gains. Rigorous convergence analysis is provided, supported by numerical experiments. The numerical experiments show excellent empirical performance of the proposed scheme.
DOI:
--
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
作者:
Akshay Mete;Rahul Singh;Xi Liu;P. Kumar
通讯作者:
Akshay Mete;Rahul Singh;Xi Liu;P. Kumar
DOI:
10.1287/opre.1070.0505
发表时间:
2009-03
期刊:
Oper. Res.
影响因子:
--
作者:
T. Archibald;Dan Black;K. Glazebrook
通讯作者:
T. Archibald;Dan Black;K. Glazebrook
影响因子:
2.7
作者:
Glazebrook K
通讯作者:
Glazebrook K