Achieving efficient interpretability of reinforcement learning via policy distillation and selective input gradient regularization

Achieving efficient interpretability of reinforcement learning via policy distillation and selective input gradient regularization
复制标题

DOI:
10.1016/j.neunet.2023.01.025
复制
发表时间:
2023-01
期刊:
Neural networks : the official journal of the International Neural Network Society
影响因子:
--
通讯作者:
Jinwei Xing;Takashi Nagata;Xinyun Zou;E. Neftci;J. Krichmar
Jinwei Xing;Takashi Nagata;Xinyun Zou;E. Neftci;J. Krichmar
中科院分区:
其他
文献类型:
--
作者:
Jinwei Xing;Takashi Nagata;Xinyun Zou;E. Neftci;J. Krichmar

文献摘要

被引文献

相似文献

尽管深度强化学习(RL)已被证明在广泛的任务中取得了成功,但它面临的一个挑战是应用于现实世界问题时的可解释性。显著性图经常被用来为深度神经网络提供可解释性。然而,在强化学习领域,现有的显著性图方法要么计算成本高,因此无法满足现实场景的实时性要求,要么无法为强化学习策略生成可解释的显著性图。在这项工作中,我们提出了一种具有选择性输入梯度正则化(DIGR)的蒸馏方法,该方法使用策略蒸馏和输入梯度正则化来生成新策略,在生成显著性映射时实现高可解释性和高计算效率。我们的方法还被发现可以提高RL策略对多个对抗性攻击的鲁棒性。我们在MiniGrid (Fetch Object)、Atari (Breakout)和CARLA Autonomous Driving三个任务上进行了实验,以证明我们的方法的重要性和有效性。
Although deep Reinforcement Learning (RL) has proven successful in a wide range of tasks, one challenge it faces is interpretability when applied to real-world problems. Saliency maps are frequently used to provide interpretability for deep neural networks. However, in the RL domain, existing saliency map approaches are either computationally expensive and thus cannot satisfy the real-time requirement of real-world scenarios or cannot produce interpretable saliency maps for RL policies. In this work, we propose an approach of Distillation with selective Input Gradient Regularization (DIGR) which uses policy distillation and input gradient regularization to produce new policies that achieve both high interpretability and computation efficiency in generating saliency maps. Our approach is also found to improve the robustness of RL policies to multiple adversarial attacks. We conduct experiments on three tasks, MiniGrid (Fetch Object), Atari (Breakout) and CARLA Autonomous Driving, to demonstrate the importance and effectiveness of our approach.