Collaborative Thompson Sampling

Collaborative Thompson Sampling
复制标题

协作汤普森采样

DOI:
10.1007/s11036-019-01453-x
复制
发表时间:
2018
影响因子:
3.8
通讯作者:
Hongli Xu
Hongli Xu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Zhenyu Zhu;Liusheng Huang;Hongli Xu

文献摘要

被引文献

相似文献

汤普森抽样是平衡勘探开发权衡的最有效策略之一。它已被应用于各种领域,并取得了显着的成功。汤普森采样通过随时间积累不确定信息来提高预测精度,从而在嘈杂但稳定的环境中做出决策。然而,在高度动态的领域中,环境经历频繁和不可预测的变化。在这样的环境中做出决定应该依赖于当前的信息。因此,标准的汤普森采样可能在这些领域表现不佳。在这里,我们提出了协作汤普森采样应用的探索,开发策略,高度动态的设置。该算法通过动态地将用户分组来考虑协作效应,并利用同一组中所有用户的反馈信息来估计当前环境下的期望回报,从而找到最优选择。将协同效应引入到汤普森采样中,可以捕捉环境的实时变化,并相应地调整决策策略。我们比较我们的算法与标准的汤普森采样算法在两个真实世界的数据集。我们的算法在协作环境中显示出加速收敛和改进的预测性能。我们还提供了我们的算法在上下文和非上下文设置的遗憾分析。
Thompson sampling is one of the most effective strategies to balance exploration-exploitation trade-off. It has been applied in a variety of domains and achieved remarkable success. Thompson sampling makes decisions in a noisy but stationary environment by accumulating uncertain information over time to improve prediction accuracy. In highly dynamic domains, however, the environment undergoes frequent and unpredictable changes. Making decisions in such an environment should rely on current information. Therefore, standard Thompson sampling may perform poorly in these domains. Here we present collaborative Thompson sampling to apply the exploration-exploitation strategy to highly dynamic settings. The algorithm takes collaborative effects into account by dynamically clustering users into groups, and the feedback of all users in the same group will help to estimate the expected reward in the current context to find the optimal choice. Incorporating collaborative effects into Thompson sampling allows to capture real-time changes of the environment and adjust decision making strategy accordingly. We compare our algorithm with standard Thompson sampling algorithms on two real-world datasets. Our algorithm shows accelerated convergence and improved prediction performance in collaborative environments. We also provide regret analyses of our algorithm in both contextual and non-contextual settings.