Gaussian Process Bandits with Aggregated Feedback

Gaussian Process Bandits with Aggregated Feedback
复制标题

具有聚合反馈的高斯过程强盗

DOI:
10.1609/aaai.v36i8.20892
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
Cheng Soon Ong
Cheng Soon Ong
中科院分区:
--
文献类型:
--
作者:
Mengyan Zhang;Russell Tsuchida;Cheng Soon Ong

文献摘要

被引文献

相似文献

我们在一种新颖的设置下考虑连续武装强盗问题,即根据汇总反馈在固定预算内推荐最佳武器。这是由应用程序推动的,在这些应用程序中获得精确奖励是不可能或昂贵的,而聚合奖励或反馈(例如子集的平均值)是可用的。我们通过假设奖励函数来自高斯过程来约束奖励函数集,并提出高斯过程乐观优化(GPOO)算法。我们自适应地构建一棵树,其中节点作为臂空间的子集,其中反馈是节点代表的聚合奖励。我们针对推荐武器的汇总反馈提出了一个新的简单遗憾概念。我们对所提出的算法进行了理论分析,并作为特例恢复了单点反馈。我们对 GPOO 进行了说明,并在模拟数据上将其与相关算法进行了比较。
We consider the continuum-armed bandits problem, under a novel setting of recommending the best arms within a fixed budget under aggregated feedback. This is motivated by applications where the precise rewards are impossible or expensive to obtain, while an aggregated reward or feedback, such as the average over a subset, is available. We constrain the set of reward functions by assuming that they are from a Gaussian Process and propose the Gaussian Process Optimistic Optimisation (GPOO) algorithm. We adaptively construct a tree with nodes as subsets of the arm space, where the feedback is the aggregated reward of representatives of a node. We propose a new simple regret notion with respect to aggregated feedback on the recommended arms. We provide theoretical analysis for the proposed algorithm, and recover single point feedback as a special case. We illustrate GPOO and compare it with related algorithms on simulated data.