Distributed Off-Policy Temporal Difference Learning Using Primal-Dual Method

Distributed Off-Policy Temporal Difference Learning Using Primal-Dual Method
复制标题

基于原对偶方法的分布式离策略时间差分学习

DOI:
10.1109/access.2022.3211395
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Donghwan Lee;Do Wan Kim;Jianghai Hu
Donghwan Lee;Do Wan Kim;Jianghai Hu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Donghwan Lee;Do Wan Kim;Jianghai Hu

文献摘要

相似文献

本文的目的是通过鞍点观点为多智能体马尔可夫决策过程(MDP)的分布式时差(TD)学习算法提供理论分析和更多的见解。(单代理)TD-学习是一种基于奖励反馈评估给定策略的强化学习(RL)算法。在多代理设置中,多个RL代理并发行为,每个代理获得其本地奖励。每个智能体的目标是评估与全局奖励相对应的给定策略,全局奖励是通过随机网络通信共享学习参数而获得的局部奖励的平均值。本文提出了一种基于鞍点框架的分布式TD-学习算法,并基于最优化理论中的工具对该算法及其解的有限时间收敛进行了严格的分析。本文的研究结果为分布式策略评估问题提供了统一的视角,在理论上补充了前人的工作。
The goal of this paper is to provide theoretical analysis and additional insights on a distributed temporal-difference (TD)-learning algorithm for the multi-agent Markov decision processes (MDPs) via saddle-point viewpoints. The (single-agent) TD-learning is a reinforcement learning (RL) algorithm for evaluating a given policy based on reward feedbacks. In multi-agent settings, multiple RL agents concurrently behave, and each agent receives its local rewards. The goal of each agent is to evaluate a given policy corresponding to the global reward, which is an average of the local rewards by sharing learning parameters through random network communications. In this paper, we propose a distributed TD-learning based on saddle-point frameworks, and provide rigorous analysis of finite-time convergence of the algorithm and its solution based on tools in optimization theory. The results in this paper provide general and unified perspectives of the distributed policy evaluation problem, and theoretically complement the previous works.