Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL

Rethinking Adversarial Policies: A Generalized Attack Formulation and Provable Defense in RL
复制标题

DOI:
--
复制
发表时间:
2023-05
期刊:
--
影响因子:
--
通讯作者:
Xiangyu Liu;Souradip Chakraborty;Yanchao Sun;Furong Huang
Xiangyu Liu;Souradip Chakraborty;Yanchao Sun;Furong Huang
中科院分区:
其他
文献类型:
--
作者:
Xiangyu Liu;Souradip Chakraborty;Yanchao Sun;Furong Huang

文献摘要

相似文献

大多数现有的作品集中在直接扰动受害者的状态/动作或潜在的过渡动态,以证明强化学习代理对抗攻击的脆弱性。然而,这种直接操纵可能并不总是可实现的。在本文中,我们考虑一个多代理设置,一个训练有素的受害者代理$\nu$被攻击者利用控制另一个代理$\alpha$与\textit{对抗策略}。以前的模型没有考虑攻击者可能只对$\alpha$有部分控制,或者攻击可能产生容易检测到的“异常”行为。此外,缺乏可证明有效的防御措施来对抗这些对抗性政策。为了解决这些限制,我们引入了一个广义的攻击框架,该框架具有灵活性,可以在多大程度上模拟对手能够控制代理,并允许攻击者调节状态分布的变化,并产生更隐蔽的对抗策略。此外,我们通过时间尺度分离的对抗性训练,提供了一个可证明有效的防御,多项式收敛到最强大的受害者策略。这与监督学习形成鲜明对比,在监督学习中,对抗训练通常只提供经验防御。使用Robosumo竞争实验,我们表明,当保持与基线相同的胜率时,我们的广义攻击公式会导致更隐蔽的对抗策略。此外,我们的对抗性训练方法产生稳定的学习动态和更少的可利用的受害者策略。
Most existing works focus on direct perturbations to the victim's state/action or the underlying transition dynamics to demonstrate the vulnerability of reinforcement learning agents to adversarial attacks. However, such direct manipulations may not be always realizable. In this paper, we consider a multi-agent setting where a well-trained victim agent $\nu$ is exploited by an attacker controlling another agent $\alpha$ with an \textit{adversarial policy}. Previous models do not account for the possibility that the attacker may only have partial control over $\alpha$ or that the attack may produce easily detectable"abnormal"behaviors. Furthermore, there is a lack of provably efficient defenses against these adversarial policies. To address these limitations, we introduce a generalized attack framework that has the flexibility to model to what extent the adversary is able to control the agent, and allows the attacker to regulate the state distribution shift and produce stealthier adversarial policies. Moreover, we offer a provably efficient defense with polynomial convergence to the most robust victim policy through adversarial training with timescale separation. This stands in sharp contrast to supervised learning, where adversarial training typically provides only \textit{empirical} defenses. Using the Robosumo competition experiments, we show that our generalized attack formulation results in much stealthier adversarial policies when maintaining the same winning rate as baselines. Additionally, our adversarial training approach yields stable learning dynamics and less exploitable victim policies.