Robust Stochastic Bandit Algorithms under Probabilistic Unbounded Adversarial Attack

Robust Stochastic Bandit Algorithms under Probabilistic Unbounded Adversarial Attack
复制标题

概率无界对抗攻击下的鲁棒随机强盗算法

DOI:
10.1609/aaai.v34i04.5821
复制
发表时间:
2020
期刊:
ArXiv
影响因子:
--
通讯作者:
Yingbin Liang
Yingbin Liang
中科院分区:
--
文献类型:
--
作者:
Ziwei Guan;Kaiyi Ji;Donald J. Bucci;Timothy Y. Hu;J. Palombo;Michael J. Liston;Yingbin Liang

文献摘要

参考文献

被引文献

相似文献

在各种攻击模型下,多武器的匪徒格式化已广泛研究,在这种模型中,对手可以修改向玩家揭示的奖励。 。每个回合的概率及其攻击值可以是任意的,如果它的攻击值不一定,则攻击值不一定遵循统计分布。 - 核心)和基于中位数的ϵ-greedy算法(称为med-greedy)。两种算法都可以实现O(log t)伪regret(即,没有攻击的最佳遗憾)。在任意和无界的奖励扰动下,只要攻击概率不超过一定的恒定阈值。展示现有技术无法实现sublerear后悔的能力。
The multi-armed bandit formalism has been extensively studied under various attack models, in which an adversary can modify the reward revealed to the player. Previous studies focused on scenarios where the attack value either is bounded at each round or has a vanishing probability of occurrence. These models do not capture powerful adversaries that can catastrophically perturb the revealed reward. This paper investigates the attack model where an adversary attacks with a certain probability at each round, and its attack value can be arbitrary and unbounded if it attacks. Furthermore, the attack value does not necessarily follow a statistical distribution. We propose a novel sample median-based and exploration-aided UCB algorithm (called med-E-UCB) and a median-based ϵ-greedy algorithm (called med-ϵ-greedy). Both of these algorithms are provably robust to the aforementioned attack model. More specifically we show that both algorithms achieve O(log T) pseudo-regret (i.e., the optimal regret without attacks). We also provide a high probability guarantee of O(log T) regret with respect to random rewards and random occurrence of attacks. These bounds are achieved under arbitrary and unbounded reward perturbation as long as the attack probability does not exceed a certain constant threshold. We provide multiple synthetic simulations of the proposed algorithms to verify these claims and showcase the inability of existing techniques to achieve sublinear regret. We also provide experimental results of the algorithm operating in a cognitive radio setting using multiple software-defined radios.
DOI: 10.1109/tit.2018.2847695
发表时间: 2016-03
影响因子: 2.5
作者:
Huishuai Zhang;Yuejie Chi;Yingbin Liang
通讯作者: Huishuai Zhang;Yuejie Chi;Yingbin Liang