Regression Oracles and Exploration Strategies for Short-Horizon Multi-Armed Bandits

Regression Oracles and Exploration Strategies for Short-Horizon Multi-Armed Bandits
复制标题

DOI:
10.1109/cog47356.2020.9231529
复制
发表时间:
2020-08
期刊:
2020 IEEE Conference on Games (CoG)
影响因子:
--
通讯作者:
Robert C. Gray;Jichen Zhu;Santiago Ontañón
Robert C. Gray;Jichen Zhu;Santiago Ontañón
中科院分区:
其他
文献类型:
--
作者:
Robert C. Gray;Jichen Zhu;Santiago Ontañón

文献摘要

被引文献

相似文献

本文探讨了非常短视野场景下的多臂老虎机(MAB)策略,即老虎机策略只允许与环境进行很少的交互。这是 MAB 文献中尚未得到充分研究的设置,在游戏环境中具有许多应用,例如玩家建模。具体来说,我们追求三种不同的想法。首先,我们探索回归预言的使用,它用线性回归模型代替 ϵ-greedy 等策略中使用的简单平均值。其次,我们研究不同的探索模式,例如强制探索阶段。最后,我们引入了 UCB1 策略的一个新变体,称为 UCBT,它具有有趣的属性并且没有可调参数。我们在运动游戏的推动下展示了实验结果,其目标是最大化玩家的日常步数。我们的结果表明,ε-贪婪或 ε-递减与回归预言的组合在短期设置中优于所有其他测试策略。
This paper explores multi-armed bandit (MAB) strategies in very short horizon scenarios, i.e., when the bandit strategy is only allowed very few interactions with the environment. This is an understudied setting in the MAB literature with many applications in the context of games, such as player modeling. Specifically, we pursue three different ideas. First, we explore the use of regression oracles, which replace the simple average used in strategies such as ϵ-greedy with linear regression models. Second, we examine different exploration patterns such as forced exploration phases. Finally, we introduce a new variant of the UCB1 strategy called UCBT that has interesting properties and no tunable parameters. We present experimental results in a domain motivated by exergames, where the goal is to maximize a player’s daily steps. Our results show that the combination of ϵ-greedy or ϵ-decreasing with regression oracles outperforms all other tested strategies in the short horizon setting.