Batched Multi-Armed Bandits with Optimal Regret

Batched Multi-Armed Bandits with Optimal Regret
复制标题

批量多臂强盗,最佳后悔

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
V. Mirrokni
V. Mirrokni
中科院分区:
--
文献类型:
--
作者:
Hossein Esfandiari;Amin Karbasi;Abbas Mehrabian;V. Mirrokni

文献摘要

参考文献

被引文献

相似文献

我们提出了一种简单有效的算法,用于批处理的随机多臂匪徒问题。我们证明了其预期的遗憾,即对任何数量的批次都改善了最著名的遗憾。特别是,我们的算法仅使用对数批次数量,从而实现了最佳的预期遗憾。
We present a simple and efficient algorithm for the batched stochastic multi-armed bandit problem. We prove a bound for its expected regret that improves over the best-known regret bound, for any number of batches. In particular, our algorithm achieves the optimal expected regret by using only a logarithmic number of batches.
DOI: --
发表时间: 2019-11
期刊: ArXiv
影响因子: --
作者:
Hossein Esfandiari;Amin Karbasi;V. Mirrokni
通讯作者: Hossein Esfandiari;Amin Karbasi;V. Mirrokni