Minimax Optimal Algorithms for Adversarial Bandit Problem With Multiple Plays
Minimax Optimal Algorithms for Adversarial Bandit Problem With Multiple Plays
复制标题
多次游戏对抗性强盗问题的最小最大最优算法
DOI:
10.1109/tsp.2019.2928952
复制
发表时间:
2019
影响因子:
5.4
通讯作者:
S. Kozat
中科院分区:
文献类型:
--
作者:
Nuri Mert Vural;Hakan Gokcesu;Kaan Gokcesu;S. Kozat
We investigate the adversarial bandit problem with multiple plays under semi-bandit feedback. We introduce a highly efficient algorithm that asymptotically achieves the performance of the best switching <inline-formula><tex-math notation="LaTeX">$m$</tex-math></inline-formula>-arm strategy with minimax optimal regret bounds. To construct our algorithm, we introduce a new expert advice algorithm for the multiple-play setting. By using our expert advice algorithm, we additionally improve the best-known high-probability bound for the multi-play setting by <inline-formula><tex-math notation="LaTeX">$O(\sqrt{m})$</tex-math></inline-formula>. Our results are guaranteed to hold in an individual sequence manner since we have no statistical assumption on the bandit arm gains. Through an extensive set of experiments involving synthetic and real data, we demonstrate significant performance gains achieved by the proposed algorithm with respect to the state-of-the-art algorithms.