The Blessing of Heterogeneity in Federated Q-learning: Linear Speedup and Beyond

The Blessing of Heterogeneity in Federated Q-learning: Linear Speedup and Beyond
复制标题

DOI:
10.48550/arxiv.2305.10697
复制
发表时间:
2023-05
期刊:
--
影响因子:
--
通讯作者:
Jiin Woo;Gauri Joshi;Yuejie Chi
Jiin Woo;Gauri Joshi;Yuejie Chi
中科院分区:
其他
文献类型:
--
作者:
Jiin Woo;Gauri Joshi;Yuejie Chi

文献摘要

相似文献

当用于强化学习(RL)的数据由多个代理以分布式方式收集时,RL算法的联邦版本允许协作学习,而无需代理共享其本地数据。在本文中,我们考虑联邦Q学习,其目的是通过定期聚合仅在局部数据上训练的局部Q估计来学习最优Q函数。专注于无限地平线表马尔可夫决策过程,我们提供了样本复杂性保证联邦Q学习的同步和异步变体。在这两种情况下,我们的边界表现出线性加速相对于代理的数量和接近最佳的依赖于其他突出的问题参数。在异步环境中,现有的联邦Q学习分析,采用了本地Q估计的平均加权,要求每个代理覆盖整个状态-动作空间。相比之下,我们改进的样本复杂度与所有代理的平均静态状态-动作占用分布的最小条目成反比,因此只需要代理共同覆盖整个状态-动作空间,通过放松单代理情况下的覆盖要求,揭示了异质性在实现协作学习中的好处。然而,当局部轨迹高度异质时,其样本复杂度仍然受到影响。作为回应,我们提出了一种新的联邦Q学习算法与重要性平均,给予更大的权重,更频繁访问的状态动作对,实现了强大的线性加速,好像所有的轨迹集中处理,无论本地行为策略的异质性。
When the data used for reinforcement learning (RL) are collected by multiple agents in a distributed manner, federated versions of RL algorithms allow collaborative learning without the need for agents to share their local data. In this paper, we consider federated Q-learning, which aims to learn an optimal Q-function by periodically aggregating local Q-estimates trained on local data alone. Focusing on infinite-horizon tabular Markov decision processes, we provide sample complexity guarantees for both the synchronous and asynchronous variants of federated Q-learning. In both cases, our bounds exhibit a linear speedup with respect to the number of agents and near-optimal dependencies on other salient problem parameters. In the asynchronous setting, existing analyses of federated Q-learning, which adopt an equally weighted averaging of local Q-estimates, require that every agent covers the entire state-action space. In contrast, our improved sample complexity scales inverse proportionally to the minimum entry of the average stationary state-action occupancy distribution of all agents, thus only requiring the agents to collectively cover the entire state-action space, unveiling the blessing of heterogeneity in enabling collaborative learning by relaxing the coverage requirement of the single-agent case. However, its sample complexity still suffers when the local trajectories are highly heterogeneous. In response, we propose a novel federated Q-learning algorithm with importance averaging, giving larger weights to more frequently visited state-action pairs, which achieves a robust linear speedup as if all trajectories are centrally processed, regardless of the heterogeneity of local behavior policies.