Simulation Studies of Multi-armed Bandits with Covariates (Invited Paper)
Simulation Studies of Multi-armed Bandits with Covariates (Invited Paper)
复制标题
具有协变量的多臂老虎机模拟研究(特邀论文)
DOI:
10.1109/uksim.2008.86
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
D. Hand
中科院分区:
文献类型:
--
作者:
N. Pavlidis;D. Tasoulis;D. Hand
We evaluate the performance of a number of action-selection methods on the multi-armed bandit problem with covariates. We resort to simulations because our primary concern is the speed with which the different methods identify the optimal policy, and not their asymptotic behaviour. The experimental results show that the performance of the ε-greedy methods is robust, while the interval estimation strategies achieve the fastest learning of the optimal policy. We propose a metric to quantify the difficulty of a multi-armed bandit problem with covariates and show that there is a trade-off between the satisfaction of the different performance measures.