Simulation Studies of Multi-armed Bandits with Covariates (Invited Paper)

Simulation Studies of Multi-armed Bandits with Covariates (Invited Paper)
复制标题

具有协变量的多臂老虎机模拟研究(特邀论文)

DOI:
10.1109/uksim.2008.86
复制
发表时间:
2008
期刊:
Tenth International Conference on Computer Modeling and Simulation (uksim 2008)
影响因子:
--
通讯作者:
D. Hand
D. Hand
中科院分区:
--
文献类型:
--
作者:
N. Pavlidis;D. Tasoulis;D. Hand

文献摘要

被引文献

相似文献

我们评估了多种动作选择方法在带有协变量的多臂老虎机问题上的性能。我们诉诸模拟,因为我们主要关心的是不同方法确定最优策略的速度,而不是它们的渐近行为。实验结果表明,ε-贪心方法的性能稳健,而区间估计策略实现了最优策略的最快学习。我们提出了一个度量来量化带有协变量的多臂老虎机问题的难度,并表明不同性能指标的满意度之间存在权衡。
We evaluate the performance of a number of action-selection methods on the multi-armed bandit problem with covariates. We resort to simulations because our primary concern is the speed with which the different methods identify the optimal policy, and not their asymptotic behaviour. The experimental results show that the performance of the ε-greedy methods is robust, while the interval estimation strategies achieve the fastest learning of the optimal policy. We propose a metric to quantify the difficulty of a multi-armed bandit problem with covariates and show that there is a trade-off between the satisfaction of the different performance measures.