Data Poisoning Attacks in Contextual Bandits

Data Poisoning Attacks in Contextual Bandits
复制标题

DOI:
10.1007/978-3-030-01554-1_11
复制
发表时间:
2018-08
期刊:
--
影响因子:
--
通讯作者:
Yuzhe Ma-;Kwang-Sung Jun;Lihong Li;Xiaojin Zhu
Yuzhe Ma-;Kwang-Sung Jun;Lihong Li;Xiaojin Zhu
中科院分区:
其他
文献类型:
--
作者:
Yuzhe Ma-;Kwang-Sung Jun;Lihong Li;Xiaojin Zhu

文献摘要

被引文献

相似文献

我们研究了上下文强盗中的离线数据中毒攻击,上下文强盗是一类强化学习问题,在在线推荐和自适应医疗等方面具有重要应用。我们提供了一个基于凸优化的通用攻击框架,并表明通过稍微操纵数据中的奖励,攻击者可以迫使强盗算法为目标上下文向量拉目标手臂。目标臂和目标上下文向量都由攻击者选择。也就是说,攻击者可以劫持上下文强盗的行为。我们还研究了这种攻击的可行性和副作用,并确定了未来的防御方向。在合成数据和真实数据上的实验证明了该攻击算法的有效性。
We study offline data poisoning attacks in contextual bandits, a class of reinforcement learning problems with important applications in online recommendation and adaptive medical treatment, among others. We provide a general attack framework based on convex optimization and show that by slightly manipulating rewards in the data, an attacker can force the bandit algorithm to pull a target arm for a target contextual vector. The target arm and target contextual vector are both chosen by the attacker. That is, the attacker can hijack the behavior of a contextual bandit. We also investigate the feasibility and the side effects of such attacks, and identify future directions for defense. Experiments on both synthetic and real-world data demonstrate the efficiency of the attack algorithm.