Model-free Learning with Heterogeneous Dynamical Systems: A Federated LQR Approach

Model-free Learning with Heterogeneous Dynamical Systems: A Federated LQR Approach
复制标题

DOI:
--
复制
发表时间:
2023-08
期刊:
--
影响因子:
--
通讯作者:
Hang Wang;Leonardo F. Toso;A. Mitra;James Anderson
Hang Wang;Leonardo F. Toso;A. Mitra;James Anderson
中科院分区:
其他
文献类型:
--
作者:
Hang Wang;Leonardo F. Toso;A. Mitra;James Anderson

文献摘要

相似文献

我们研究了一个无模型的联合线性二次调节器(LQR)问题,其中的M代理商具有未知,不同但相似的动态的代理协作,可以协作学习最佳政策,以最大程度地减少平均二次成本,同时保持其数据私密。为了利用代理动力学的相似性,我们建议使用联合学习(FL),以允许代理商与中央服务器定期通信以通过利用所有代理商的较大数据集来培训策略。有了这种设置,我们试图理解以下问题:(i)对所有代理商来说,学到的共同政策是否稳定? (ii)学到的公共政策与每个代理自己的最佳政策有多近? (iii)每个代理可以通过利用所有代理的数据来更快地学习自己的最佳政策?为了回答这些问题,我们提出了一种名为FedLQR的联合且无模型的算法。我们的分析克服了许多技术挑战,例如代理商的动态,多个本地更新和稳定问题的异质性。我们表明,FedLQR制定了一项共同的政策,在每次迭代中,所有代理商都在稳定。我们提供有关共同政策与每个代理人本地最佳政策之间距离的界限。此外,我们证明,在学习每个代理的最佳策略时,FedLQR与单代机构设置相比,FedLQR实现了与低杂种性制度中M数量成正比的样本复杂性降低。
We study a model-free federated linear quadratic regulator (LQR) problem where M agents with unknown, distinct yet similar dynamics collaboratively learn an optimal policy to minimize an average quadratic cost while keeping their data private. To exploit the similarity of the agents' dynamics, we propose to use federated learning (FL) to allow the agents to periodically communicate with a central server to train policies by leveraging a larger dataset from all the agents. With this setup, we seek to understand the following questions: (i) Is the learned common policy stabilizing for all agents? (ii) How close is the learned common policy to each agent's own optimal policy? (iii) Can each agent learn its own optimal policy faster by leveraging data from all agents? To answer these questions, we propose a federated and model-free algorithm named FedLQR. Our analysis overcomes numerous technical challenges, such as heterogeneity in the agents' dynamics, multiple local updates, and stability concerns. We show that FedLQR produces a common policy that, at each iteration, is stabilizing for all agents. We provide bounds on the distance between the common policy and each agent's local optimal policy. Furthermore, we prove that when learning each agent's optimal policy, FedLQR achieves a sample complexity reduction proportional to the number of agents M in a low-heterogeneity regime, compared to the single-agent setting.