Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption

Restless Bandits with Average Reward: Breaking the Uniform Global Attractor Assumption
复制标题

平均奖励的不安分强盗:打破统一的全球吸引子假设

DOI:
--
复制
发表时间:
2023
期刊:
Advances in neural information processing systems
影响因子:
--
通讯作者:
Wang, Weina
Wang, Weina
中科院分区:
--
文献类型:
--
作者:
Hong, Yige;Xie, Qiaomin;Chen, Yudong;Wang, Weina

文献摘要

参考文献

被引文献

相似文献

DOI: 10.1287/opre.1070.0505
发表时间: 2009-03
期刊: Oper. Res.
影响因子: --
作者:
T. Archibald;Dan Black;K. Glazebrook
通讯作者: T. Archibald;Dan Black;K. Glazebrook
动态选择问题的索引策略和性能界限
DOI: --
发表时间: 2020
期刊: Management Sciences
影响因子: --
作者:
David B. Brown;James E. Smith
通讯作者: James E. Smith
对于 Whittle 来说,可索引性还不够:针对不安分强盗的改进的、近乎最优的算法
DOI: --
发表时间: 2022
期刊: Adaptive Agents and Multi-Agent Systems
影响因子: --
作者:
Abheek Ghosh;Dheeraj M. Nagaraj;Manish Jain;Milind Tambe
通讯作者: Milind Tambe
多臂不安的强盗:击败中心极限定理
DOI: --
发表时间: 2021
期刊: arXiv.org
影响因子: --
作者:
X. Zhang;P. Frazier
通讯作者: P. Frazier
一般非平稳有限视野不安定多臂多动作老虎机的渐近最优启发式
DOI: --
发表时间: 2017
影响因子: 1.2
作者:
Gabriel Zayas;Stefanus Jasin;Guihua Wang
通讯作者: Guihua Wang