Policy Smoothing for Provably Robust Reinforcement Learning

Policy Smoothing for Provably Robust Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Aounon Kumar;Alexander Levine;S. Feizi
Aounon Kumar;Alexander Levine;S. Feizi
中科院分区:
其他
文献类型:
--
作者:
Aounon Kumar;Alexander Levine;S. Feizi

文献摘要

相似文献

深度神经网络(DNN)的可证明对抗鲁棒性研究主要集中在静态监督学习任务,如图像分类。然而,DNN已广泛用于现实世界的自适应任务,例如强化学习(RL),这使得此类系统也容易受到对抗性攻击。以前的工作在RL的可证明的鲁棒性,试图证明受害者政策的行为在每个时间步对非自适应对手使用的方法开发的静态设置。但在真实的世界中,RL对手可以通过观察来自先前时间步的状态、动作等来推断受害者代理所使用的防御策略,并调整自己以在未来步骤中产生更强的攻击(例如,通过更多地关注对代理的性能至关重要的状态)。我们提出了一个有效的程序,专门设计用于防御自适应RL对手,可以直接证明总奖励,而不需要在每个时间步都保持策略的鲁棒性。专注于基于随机平滑的防御,我们的主要理论贡献是证明了Neyman-Pearson引理的自适应版本-基于平滑的证书的关键引理-其中特定时间的对抗性扰动可以是当前和先前观测和状态以及先前动作的随机函数。基于这一结果,我们提出了策略平滑,其中代理在每个时间步将高斯噪声添加到其观察中,然后将其传递给策略函数。我们的鲁棒性保证了通过策略平滑获得的最终总回报保持在一定的阈值以上,即使中间时间步的动作可能会在攻击下发生变化。我们通过构建一个最坏情况的场景来证明我们的证书是紧的,该场景达到了我们分析中得出的界限。我们在Cartpole、Pong、Freeway和Mountain Car等各种环境下的实验表明,我们的方法可以在实践中产生有意义的鲁棒性保证。
The study of provable adversarial robustness for deep neural networks (DNNs) has mainly focused on static supervised learning tasks such as image classification. However, DNNs have been used extensively in real-world adaptive tasks such as reinforcement learning (RL), making such systems vulnerable to adversarial attacks as well. Prior works in provable robustness in RL seek to certify the behaviour of the victim policy at every time-step against a non-adaptive adversary using methods developed for the static setting. But in the real world, an RL adversary can infer the defense strategy used by the victim agent by observ-ing the states, actions, etc. from previous time-steps and adapt itself to produce stronger attacks in future steps (e.g., by focusing more on states critical to the agent’s performance). We present an efficient procedure, designed specifically to defend against an adaptive RL adversary, that can directly certify the total reward without requiring the policy to be robust at each time-step. Focusing on randomized smoothing based defenses, our main theoretical contribution is to prove an adaptive version of the Neyman-Pearson Lemma – a key lemma for smoothing-based certificates – where the adversarial perturbation at a particular time can be a stochastic function of current and previous observations and states as well as previous actions. Building on this result, we propose policy smoothing where the agent adds a Gaussian noise to its observation at each time-step before passing it through the policy function. Our robustness certificates guarantee that the final total reward obtained by policy smoothing remains above a certain threshold, even though the actions at intermediate time-steps may change under the attack. We show that our certificates are tight by constructing a worst-case scenario that achieves the bounds derived in our analysis. Our experiments on various environments like Cartpole, Pong, Freeway and Mountain Car show that our method can yield meaningful robustness guarantees in practice.