Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled Perturbations

Game-Theoretic Robust Reinforcement Learning Handles Temporally-Coupled Perturbations
复制标题

DOI:
10.48550/arxiv.2307.12062
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
Yongyuan Liang;Yanchao Sun;Ruijie Zheng;Xiangyu Liu;T. Sandholm;Furong Huang;S. McAleer
Yongyuan Liang;Yanchao Sun;Ruijie Zheng;Xiangyu Liu;T. Sandholm;Furong Huang;S. McAleer
中科院分区:
其他
文献类型:
--
作者:
Yongyuan Liang;Yanchao Sun;Ruijie Zheng;Xiangyu Liu;T. Sandholm;Furong Huang;S. McAleer

文献摘要

被引文献

相似文献

鲁棒强化学习(RL)旨在训练能够在环境扰动或对抗性攻击下表现良好的策略。现有方法通常假设可能的扰动空间在时间步长上保持相同。然而,在许多设置中,给定时间步长可能的扰动空间取决于过去的扰动。我们正式引入了时间耦合扰动,为现有的鲁棒强化学习方法提出了新的挑战。为了应对这一挑战,我们提出了 GRAD,这是一种新颖的博弈论方法,它将时间耦合的鲁棒强化学习问题视为部分可观察的两人零和博弈。通过在这个博弈中找到近似平衡,GRAD 确保了智能体对抗时间耦合扰动的鲁棒性。对各种连续控制任务的实证实验表明,与针对标准和时间耦合攻击的基线相比,我们提出的方法在状态和动作空间中都表现出显着的鲁棒性优势。
Robust reinforcement learning (RL) seeks to train policies that can perform well under environment perturbations or adversarial attacks. Existing approaches typically assume that the space of possible perturbations remains the same across timesteps. However, in many settings, the space of possible perturbations at a given timestep depends on past perturbations. We formally introduce temporally-coupled perturbations, presenting a novel challenge for existing robust RL methods. To tackle this challenge, we propose GRAD, a novel game-theoretic approach that treats the temporally-coupled robust RL problem as a partially-observable two-player zero-sum game. By finding an approximate equilibrium in this game, GRAD ensures the agent’s robustness against temporally-coupled perturbations. Empirical experiments on a variety of continuous control tasks demonstrate that our proposed approach exhibits significant robustness advantages compared to baselines against both standard and temporally-coupled attacks, in both state and action spaces.