Enforcing Signal Temporal Logic Specifications in Multi-Agent Adversarial Environments: A Deep Q-Learning Approach
Enforcing Signal Temporal Logic Specifications in Multi-Agent Adversarial Environments: A Deep Q-Learning Approach
复制标题
在多智能体对抗环境中执行信号时态逻辑规范:深度 Q 学习方法
DOI:
10.1109/cdc.2018.8618746
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Farhood, M.
中科院分区:
文献类型:
--
作者:
Muniraj, D;Vamvoudakis, K. G;Farhood, M.
This work addresses the problem of learning optimal control policies for a multi-agent system in an adversarial environment. Specifically, we focus on multi-agent systems where the mission objectives are expressed as signal temporal logic (STL) specifications. The agents are classified as either defensive or adversarial. The defensive agents are maximizers, namely, they maximize an objective function that enforces the STL specification; the adversarial agents, on the other hand, are minimizers. The interaction among the agents is modeled as a finite-state team stochastic game with an unknown transition probability function. The synthesis objective is to determine optimal control policies for the defensive agents that implement the STL specification against the best responses of the adversarial agents. A multi-agent deep Q-learning algorithm, which is an extension of the minimax Q-learning algorithm, is then proposed to learn the optimal policies. The effectiveness of the proposed approach is illustrated through a simulation case study.
DOI:
--
发表时间:
2002
期刊:
Neural Information Processing Systems
影响因子:
--
作者:
M. Lagoudakis;Ronald E. Parr
通讯作者:
Ronald E. Parr
DOI:
--
发表时间:
2012
期刊:
Time
影响因子:
--
作者:
C. Belta
通讯作者:
C. Belta