ACADIA: Efficient and Robust Adversarial Attacks Against Deep Reinforcement Learning

ACADIA: Efficient and Robust Adversarial Attacks Against Deep Reinforcement Learning
复制标题

DOI:
10.1109/cns56114.2022.9947234
复制
发表时间:
2022-10
期刊:
2022 IEEE Conference on Communications and Network Security (CNS)
影响因子:
--
通讯作者:
Haider Ali;Mohannad Al Ameedi;A. Swami;R. Ning;Jiang Li;Hongyi Wu;Jin-Hee Cho
Haider Ali;Mohannad Al Ameedi;A. Swami;R. Ning;Jiang Li;Hongyi Wu;Jin-Hee Cho
中科院分区:
其他
文献类型:
--
作者:
Haider Ali;Mohannad Al Ameedi;A. Swami;R. Ning;Jiang Li;Hongyi Wu;Jin-Hee Cho

文献摘要

相似文献

现有的深入加固学习(DRL)的对抗算法主要集中在确定攻击DRL代理的最佳时间。但是,在DRL环境中注入有效的对抗扰动方面,几乎没有探索工作。我们提出了一套新型的DRL对抗性攻击,称为Acadia,代表了针对深入强化学习的攻击。阿卡迪亚(Acadia)提供了一套有效且强大的基于扰动的对抗攻击,以利用动量,Adam Optimizer(即根平方敏化或RMSPROP)和初始随机化的技术组合来干扰DRL代理的决策。在现有的深层神经网络(DNNS)和DRL研究中,尚未研究这种技术的新型DRL攻击。我们在Atari Games和Mujoco下考虑了两种众所周知的DRL算法,即深Q学习网络(DQN)和近端政策优化(PPO)(PPO),其中有针对性和非目标攻击都在有或没有最新的情况下考虑。 DRL(即径向和Atla)中的艺术防御。我们的结果表明,拟议的Acadia在广泛的实验环境下优于现有的基于梯度的同行。 Acadia的速度比最先进的Carlini&Wagner(CW)方法快九倍,在DRL的防御下性能更好。
Existing adversarial algorithms for Deep Reinforcement Learning (DRL) have largely focused on identifying an optimal time to attack a DRL agent. However, little work has been explored in injecting efficient adversarial perturbations in DRL environments. We propose a suite of novel DRL adversarial attacks, called ACADIA, representing AttaCks Against Deep reInforcement leArning. ACADIA provides a set of efficient and robust perturbation-based adversarial attacks to disturb the DRL agent's decision-making based on novel combinations of techniques utilizing momentum, ADAM optimizer (i.e., Root Mean Square Propagation, or RMSProp), and initial randomization. These kinds of DRL attacks with novel integration of such techniques have not been studied in the existing Deep Neural Networks (DNNs) and DRL research. We consider two well-known DRL algorithms, Deep-Q Learning Network (DQN) and Proximal Policy Optimization (PPO), under Atari games and MuJoCo where both targeted and non-targeted attacks are considered with or without the state-of-the-art defenses in DRL (i.e., RADIAL and ATLA). Our results demonstrate that the proposed ACADIA outperforms existing gradient-based counterparts under a wide range of experimental settings. ACADIA is nine times faster than the state-of-the-art Carlini & Wagner (CW) method with better performance under defenses of DRL.