Sample Efficient Stochastic Policy Extragradient Algorithm for Zero-Sum Markov Game
Sample Efficient Stochastic Policy Extragradient Algorithm for Zero-Sum Markov Game
复制标题
DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Ziyi Chen;Shaocong Ma;Yi Zhou
中科院分区:
文献类型:
--
作者:
Ziyi Chen;Shaocong Ma;Yi Zhou
Two-player zero-sum Markov game is a fundamental problem in reinforcement learning and game theory. Although many algorithms have been proposed for solving zero-sum Markov games in the existing literature, many of them either require a full knowledge of the environment or are not sample-efficient. In this paper, we develop a fully decentralized and sample-efficient stochastic policy extragradient algorithm for solving tabular zero-sum Markov games. In particular, our algorithm utilizes multiple stochastic estimators to accurately estimate the value functions involved in the stochastic updates, and leverages entropy regularization to accelerate the convergence. Specifically, with a proper entropy-regularization parameter, we prove that the stochastic policy extragradient algorithm has a sample complexity of the order (cid:101) O ( A max µ min (cid:15) 5 . 5 (1 − γ ) 13 . 5 ) for finding a solution that achieves (cid:15) -Nash equilibrium duality gap, where A max is the maximum number of actions between the players, µ min is the lower bound of state stationary distribution, and γ is the discount factor. Such a sample complexity result substantially improves the state-of-the-art complexity result.