Learning Nash Equilibria in Zero-Sum Stochastic Games via Entropy-Regularized Policy Approximation

Learning Nash Equilibria in Zero-Sum Stochastic Games via Entropy-Regularized Policy Approximation
复制标题

DOI:
10.24963/ijcai.2021/339
复制
发表时间:
2020-09
期刊:
--
影响因子:
--
通讯作者:
Qifan Zhang;Yue Guan;P. Tsiotras
Qifan Zhang;Yue Guan;P. Tsiotras
中科院分区:
其他
文献类型:
--
作者:
Qifan Zhang;Yue Guan;P. Tsiotras

文献摘要

被引文献

相似文献

我们探讨使用政策近似,以减少零和随机博弈中学习纳什均衡的计算成本。我们提出了一种新的Q学习型算法,它使用一系列熵正则化的软策略来近似纳什策略在Q函数更新。我们证明了在一定条件下,通过更新熵正则化,算法收敛到纳什均衡。我们还证明了所提出的算法能够转移之前的训练经验,使代理能够快速适应新环境。我们提供了一个动态的超参数调度方案,以进一步加快收敛。应用于一些随机游戏的实证结果验证,该算法收敛到纳什均衡,同时表现出一个主要的速度比现有的算法。
We explore the use of policy approximations to reduce the computational cost of learning Nash equilibria in zero-sum stochastic games. We propose a new Q-learning type algorithm that uses a sequence of entropy-regularized soft policies to approximate the Nash policy during the Q-function updates. We prove that under certain conditions, by updating the entropy regularization, the algorithm converges to a Nash equilibrium. We also demonstrate the proposed algorithm's ability to transfer previous training experiences, enabling the agents to adapt quickly to new environments. We provide a dynamic hyper-parameter scheduling scheme to further expedite convergence. Empirical results applied to a number of stochastic games verify that the proposed algorithm converges to the Nash equilibrium, while exhibiting a major speed-up over existing algorithms.