Adversarial Group Linear Bandits and Its Application to Collaborative Edge Inference

Adversarial Group Linear Bandits and Its Application to Collaborative Edge Inference
复制标题

DOI:
10.1109/infocom53939.2023.10228900
复制
发表时间:
2023-05
期刊:
IEEE INFOCOM 2023 - IEEE Conference on Computer Communications
影响因子:
--
通讯作者:
Yin-Hae Huang;Letian Zhang;J. Xu
Yin-Hae Huang;Letian Zhang;J. Xu
中科院分区:
其他
文献类型:
--
作者:
Yin-Hae Huang;Letian Zhang;J. Xu

文献摘要

相似文献

多武装抢匪问题是一个经典的不确定序贯决策问题。现有的大多数作品研究土匪问题的随机奖励制度或对抗性奖励制度,但这两个制度的交叉研究少得多。在本文中,我们研究了一个新的土匪问题,称为对抗群体线性土匪(AGLB),其特点是奖励生成作为随机过程和对抗行为的联合结果。特别是,学习者获得的奖励不仅是学习者在组内选择的手臂的噪声线性函数,而且还取决于对手的组级攻击决策。这样的问题存在于许多现实世界的应用中,例如,协作边缘推断和多站点在线广告投放。为了克服随机和对抗性奖励耦合的不确定性,我们开发了一种新的强盗算法,称为EXPUCB,它结合了经典的LinUCB和EXP3算法,并证明了它的次线性遗憾。我们将EXPUCB应用于协作边缘推理问题,并评估其性能。大量的仿真结果验证了EXPUCB在耦合随机和对抗性奖励下的上级学习能力。
Multi-armed bandits is a classical sequential decision-making under uncertainty problem. The majority of existing works study bandits problems in either the stochastic reward regime or the adversarial reward regime, but the intersection of these two regimes is much less investigated. In this paper, we study a new bandits problem, called adversarial group linear bandits (AGLB), that features reward generation as a joint outcome of both the stochastic process and the adversarial behavior. In particular, the reward that the learner receives is not only a noisy linear function of the arm that the learner selects within a group but also depends on the group-level attack decision by the adversary. Such problems are present in many real-world applications, e.g., collaborative edge inference and multi-site online ad placement. To combat the uncertainty in the coupled stochastic and adversarial rewards, we develop a new bandits algorithm, called EXPUCB, which marries the classical LinUCB and EXP3 algorithms, and prove its sublinear regret. We apply EXPUCB to the collaborative edge inference problem and evaluate its performance. Extensive simulation results verify the superior learning ability of EXPUCB under coupled stochastic and adversarial rewards.