Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Stable Algorithms

Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Stable Algorithms
复制标题

DOI:
10.48550/arxiv.2310.10810
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Alexander W. Bukharin;Yan Li;Yue Yu;Qingru Zhang;Zhehui Chen;Simiao Zuo;Chao Zhang;Songan Zhang;Tuo Zhao
Alexander W. Bukharin;Yan Li;Yue Yu;Qingru Zhang;Zhehui Chen;Simiao Zuo;Chao Zhang;Songan Zhang;Tuo Zhao
中科院分区:
其他
文献类型:
--
作者:
Alexander W. Bukharin;Yan Li;Yue Yu;Qingru Zhang;Zhehui Chen;Simiao Zuo;Chao Zhang;Songan Zhang;Tuo Zhao

文献摘要

被引文献

相似文献

多代理增强学习(MARL)在几个领域显示出令人鼓舞的结果。尽管有这样的承诺,MARL政策通常缺乏稳健性,因此对环境的微小变化敏感。这对MARL算法的现实世界部署引起了人们的严重关注,其中测试环境可能与培训环境略有不同。在这项工作中,我们表明,我们可以通过控制政策的Lipschitz常数来获得鲁棒性,并且在温和的条件下,建立了Lipschitz的存在和近乎最佳的政策。基于这些见解,我们提出了一个新的强大的MARL框架Ernie,该框架促进了对逆境正则化的国家观察和行动的Lipschitz连续性。 Ernie框架为嘈杂的观察,不断变化的过渡动态和恶意行为提供了鲁棒性。但是,Ernie的对抗正规化可能会引入一些训练不稳定。为了减少这种不稳定,我们将对抗性正则化作为Stackelberg游戏。我们通过在交通灯控制和粒子环境中进行了广泛的实验来证明所提出的框架的有效性。此外,我们将Ernie扩展到平均田地MAR,其基于​​分布强大的优化的配方,以优于其非持续的对应物,并且具有独立的利益。我们的代码可在https://github.com/abukharin3/ernie上找到。
Multi-Agent Reinforcement Learning (MARL) has shown promising results across several domains. Despite this promise, MARL policies often lack robustness and are therefore sensitive to small changes in their environment. This presents a serious concern for the real world deployment of MARL algorithms, where the testing environment may slightly differ from the training environment. In this work we show that we can gain robustness by controlling a policy's Lipschitz constant, and under mild conditions, establish the existence of a Lipschitz and close-to-optimal policy. Based on these insights, we propose a new robust MARL framework, ERNIE, that promotes the Lipschitz continuity of the policies with respect to the state observations and actions by adversarial regularization. The ERNIE framework provides robustness against noisy observations, changing transition dynamics, and malicious actions of agents. However, ERNIE's adversarial regularization may introduce some training instability. To reduce this instability, we reformulate adversarial regularization as a Stackelberg game. We demonstrate the effectiveness of the proposed framework with extensive experiments in traffic light control and particle environments. In addition, we extend ERNIE to mean-field MARL with a formulation based on distributionally robust optimization that outperforms its non-robust counterpart and is of independent interest. Our code is available at https://github.com/abukharin3/ERNIE.