Multi-armed Bandit Problems with History
Multi-armed Bandit Problems with History
复制标题
历史上的多臂强盗问题
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
T. Joachims
中科院分区:
文献类型:
--
作者:
Pannagadatta K. Shivaswamy;T. Joachims
In this paper we consider the stochastic multi-armed bandit problem. However, unlike in the conventional version of this problem, we do not assume that the algorithm starts from scratch. Many applications offer observations of (some of) the arms even before the algorithm starts. We propose three novel multi-armed bandit algorithms that can exploit this data. An upper bound on the regret is derived in each case. The results show that a logarithmic amount of historic data can reduce regret from logarithmic to constant. The eectiveness of the proposed algorithms are demonstrated on a large-scale malicious URL detection problem.