Distributed Q-Learning with State Tracking for Multi-agent Networked Control

Distributed Q-Learning with State Tracking for Multi-agent Networked Control
复制标题

DOI:
10.5555/3463952.3464203
复制
发表时间:
2020-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Hang Wang;Sen Lin;H. Jafarkhani;Junshan Zhang
Hang Wang;Sen Lin;H. Jafarkhani;Junshan Zhang
中科院分区:
其他
文献类型:
--
作者:
Hang Wang;Sen Lin;H. Jafarkhani;Junshan Zhang

文献摘要

被引文献

相似文献

研究了多智能体网络中线性二次型调节器(LQR)的分布式Q学习问题。现有的结果往往假设代理可以观察到的全球系统状态,这可能是不可行的,在大规模的系统,由于隐私问题或通信约束。在这项工作中,我们考虑具有未知系统模型且没有集中协调器的设置。我们设计了一个基于状态跟踪(ST)的Q学习算法来设计最优控制器的代理。具体来说,我们假设代理维护本地估计的全球状态的基础上,他们的本地信息和与邻居的通信。在每一步中,每个代理更新其局部全局状态估计,在此基础上,它通过策略迭代在局部求解近似Q因子。假设在策略评估过程中注入衰减的激励噪声,我们证明了局部估计收敛于真正的全局状态,并建立了所提出的基于ST的分布式Q学习算法的收敛性。实验研究证实了我们的理论结果表明,我们提出的方法实现了类似的性能与集中的情况下。
This paper studies distributed Q-learning for Linear Quadratic Regulator (LQR) in a multi-agent network. The existing results often assume that agents can observe the global system state, which may be infeasible in large-scale systems due to privacy concerns or communication constraints. In this work, we consider a setting with unknown system models and no centralized coordinator. We devise a state tracking (ST) based Q-learning algorithm to design optimal controllers for agents. Specifically, we assume that agents maintain local estimates of the global state based on their local information and communications with neighbors. At each step, every agent updates its local global state estimation, based on which it solves an approximate Q-factor locally through policy iteration. Assuming decaying injected excitation noise during the policy evaluation, we prove that the local estimation converges to the true global state, and establish the convergence of the proposed distributed ST-based Q-learning algorithm. The experimental studies corroborate our theoretical results by showing that our proposed method achieves comparable performance with the centralized case.