Enforcing Signal Temporal Logic Specifications in Multi-Agent Adversarial Environments: A Deep Q-Learning Approach

Enforcing Signal Temporal Logic Specifications in Multi-Agent Adversarial Environments: A Deep Q-Learning Approach
复制标题

在多智能体对抗环境中执行信号时态逻辑规范:深度 Q 学习方法

DOI:
10.1109/cdc.2018.8618746
复制
发表时间:
2019
期刊:
2018 IEEE Conference on Decision and Control (CDC
影响因子:
--
通讯作者:
Farhood, M.
Farhood, M.
中科院分区:
--
文献类型:
--
作者:
Muniraj, D;Vamvoudakis, K. G;Farhood, M.

文献摘要

参考文献

被引文献

相似文献

这项工作解决了在对抗环境中学习多智能体系统的最优控制策略的问题。具体来说,我们专注于多智能体系统的使命目标表示为信号时序逻辑(STL)规范。这些代理人被分类为防御性或对抗性。防御性代理是最大化者,也就是说,他们最大化一个执行STL规范的目标函数;另一方面,对抗性代理是最小化者。代理之间的相互作用被建模为一个有限状态的团队随机博弈与未知的转移概率函数。合成的目标是确定最佳控制策略的防御代理,实现STL规范对敌对代理的最佳反应。一个多智能体深度Q学习算法,这是一个扩展的极小极大Q学习算法,然后提出学习的最优策略。仿真实例表明了该方法的有效性。
This work addresses the problem of learning optimal control policies for a multi-agent system in an adversarial environment. Specifically, we focus on multi-agent systems where the mission objectives are expressed as signal temporal logic (STL) specifications. The agents are classified as either defensive or adversarial. The defensive agents are maximizers, namely, they maximize an objective function that enforces the STL specification; the adversarial agents, on the other hand, are minimizers. The interaction among the agents is modeled as a finite-state team stochastic game with an unknown transition probability function. The synthesis objective is to determine optimal control policies for the defensive agents that implement the STL specification against the best responses of the adversarial agents. A multi-agent deep Q-learning algorithm, which is an extension of the minimax Q-learning algorithm, is then proposed to learn the optimal policies. The effectiveness of the proposed approach is illustrated through a simulation case study.
使用分解值函数在零和团队马尔可夫博弈中学习
DOI: --
发表时间: 2002
期刊: Neural Information Processing Systems
影响因子: --
作者:
M. Lagoudakis;Ronald E. Parr
通讯作者: Ronald E. Parr
动力系统的形式化方法
DOI: --
发表时间: 2012
期刊: Time
影响因子: --
作者:
C. Belta
通讯作者: C. Belta