An improved upper bound on the expected regret of UCB-type policies for a matching-selection bandit problem

An improved upper bound on the expected regret of UCB-type policies for a matching-selection bandit problem
复制标题

针对匹配选择强盗问题的 UCB 型策略的预期遗憾的改进上限

DOI:
10.1016/j.orl.2015.08.008
复制
发表时间:
2015
影响因子:
1.1
通讯作者:
Mineichi Kudo
Mineichi Kudo
中科院分区:
管理学4区
文献类型:
--
作者:
Ryo Watanabe;Atsuyoshi Nakamura;Mineichi Kudo

文献摘要

相似文献