Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning

Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Amin Rakhsha;Goran Radanovic;R. Devidze;Xiaojin Zhu;A. Singla
Amin Rakhsha;Goran Radanovic;R. Devidze;Xiaojin Zhu;A. Singla
中科院分区:
其他
文献类型:
--
作者:
Amin Rakhsha;Goran Radanovic;R. Devidze;Xiaojin Zhu;A. Singla

文献摘要

相似文献

我们研究强化学习的安全威胁,攻击者毒害学习环境,迫使代理执行攻击者选择的目标策略。作为受害者,我们考虑强化学习代理,其目标是找到一种策略,在未贴现的无限范围问题设置中最大化平均奖励。攻击者可以在训练时操纵学习环境中的奖励或过渡动态,并且有兴趣以隐秘的方式这样做。我们提出了一个优化框架,用于针对不同的攻击成本度量找到\emph{最佳隐形攻击}。我们提供了攻击可行的充分技术条件,并提供了攻击成本的下限/上限。我们在两种设置中实例化攻击:(i)\emph{offline} 设置,其中代理在中毒环境中进行规划;(ii)\emph{online} 设置,其中代理使用带有中毒反馈的遗憾最小化框架来学习策略。我们的结果表明,攻击者可以在温和的条件下轻松成功地向受害者教授任何目标策略,并在实践中凸显对强化学习代理的重大安全威胁。
We study a security threat to reinforcement learning where an attacker poisons the learning environment to force the agent into executing a target policy chosen by the attacker. As a victim, we consider RL agents whose objective is to find a policy that maximizes average reward in undiscounted infinite-horizon problem settings. The attacker can manipulate the rewards or the transition dynamics in the learning environment at training-time and is interested in doing so in a stealthy manner. We propose an optimization framework for finding an \emph{optimal stealthy attack} for different measures of attack cost. We provide sufficient technical conditions under which the attack is feasible and provide lower/upper bounds on the attack cost. We instantiate our attacks in two settings: (i) an \emph{offline} setting where the agent is doing planning in the poisoned environment, and (ii) an \emph{online} setting where the agent is learning a policy using a regret-minimization framework with poisoned feedback. Our results show that the attacker can easily succeed in teaching any target policy to the victim under mild conditions and highlight a significant security threat to reinforcement learning agents in practice.