Falsification-Based Robust Adversarial Reinforcement Learning

Falsification-Based Robust Adversarial Reinforcement Learning
复制标题

DOI:
10.1109/icmla51294.2020.00042
复制
发表时间:
2020-07
期刊:
2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA)
影响因子:
--
通讯作者:
Xiao Wang;Saasha Nair;M. Althoff
Xiao Wang;Saasha Nair;M. Althoff
中科院分区:
其他
文献类型:
--
作者:
Xiao Wang;Saasha Nair;M. Althoff

文献摘要

被引文献

相似文献

强化学习(RL)在解决各种顺序决策问题方面取得了巨大的进展,例如机器人中的控制任务。由于策略过于适合培训环境,因此RL方法往往无法推广到安全关键的测试场景。稳健对抗性RL(RARL)以前被提出用来训练对系统施加干扰的对抗性网络,从而提高测试场景中的稳健性。然而,基于神经网络的对手的一个问题是,在不手工制作复杂的奖励信号的情况下集成系统需求是困难的。安全伪造方法允许人们找到一组初始条件和输入序列,使得系统违反了在时序逻辑中公式化的给定属性。在本文中,我们提出了基于证伪的RARL(FRARL):这是第一个将时态逻辑证伪集成到对抗性学习中以提高策略健壮性的通用框架。通过应用我们的证伪方法,我们不需要为对手构造额外的奖励函数。此外,我们对自主车辆的制动辅助系统和自适应巡航控制系统的方法进行了评估。我们的实验结果表明,与没有对手或使用对手网络训练的策略相比,使用基于证伪的对手训练的策略具有更好的泛化能力,并且在测试场景中对安全规范的违反更少。
Reinforcement learning (RL) has achieved enormous progress in solving various sequential decision-making problems, such as control tasks in robotics. Since policies are overfitted to training environments, RL methods have often failed to be generalized to safety-critical test scenarios. Robust adversarial RL (RARL) was previously proposed to train an adversarial network that applies disturbances to a system, which improves the robustness in test scenarios. However, an issue of neural network-based adversaries is that integrating system requirements without handcrafting sophisticated reward signals are difficult. Safety falsification methods allow one to find a set of initial conditions and an input sequence, such that the system violates a given property formulated in temporal logic. In this paper, we propose falsification-based RARL (FRARL): this is the first generic framework for integrating temporal logic falsification in adversarial learning to improve policy robustness. By applying our falsification method, we do not need to construct an extra reward function for the adversary. Moreover, we evaluate our approach on a braking assistance system and an adaptive cruise control system of autonomous vehicles. Our experimental results demonstrate that policies trained with a falsification-based adversary generalize better and show less violation of the safety specification in test scenarios than those trained without an adversary or with an adversarial network.