Whittle index based Q-learning for restless bandits with average reward

Whittle index based Q-learning for restless bandits with average reward
复制标题

基于 Whittle 指数的 Q 学习,用于具有平均奖励的不安强盗

DOI:
10.1016/j.automatica.2022.110186
复制
发表时间:
2020
期刊:
ArXiv
影响因子:
--
通讯作者:
V. Borkar
V. Borkar
中科院分区:
--
文献类型:
--
作者:
Konstantin Avrachenkov;V. Borkar

文献摘要

参考文献

被引文献

相似文献

采用q -学习和Whittle指数的方法,提出了一种针对平均奖励的多臂不宁盗匪的强化学习算法。具体来说,我们利用Whittle索引策略的结构来减少Q-learning的搜索空间,从而获得重大的计算收益。在数值实验的支持下,给出了严格的收敛分析。数值实验表明,该方法具有良好的经验性能。
A novel reinforcement learning algorithm is introduced for multiarmed restless bandits with average reward, using the paradigms of Q-learning and Whittle index. Specifically, we leverage the structure of the Whittle index policy to reduce the search space of Q-learning, resulting in major computational gains. Rigorous convergence analysis is provided, supported by numerical experiments. The numerical experiments show excellent empirical performance of the proposed scheme.
DOI: --
发表时间: 2020-11
期刊: ArXiv
影响因子: --
作者:
Akshay Mete;Rahul Singh;Xi Liu;P. Kumar
通讯作者: Akshay Mete;Rahul Singh;Xi Liu;P. Kumar
DOI: 10.1287/opre.1070.0505
发表时间: 2009-03
期刊: Oper. Res.
影响因子: --
作者:
T. Archibald;Dan Black;K. Glazebrook
通讯作者: T. Archibald;Dan Black;K. Glazebrook
DOI: 10.1287/opre.1080.0632
发表时间: 2009
影响因子: 2.7
作者:
Glazebrook K
通讯作者: Glazebrook K