Model-Free Reinforcement Learning for Stochastic Parity Games

Model-Free Reinforcement Learning for Stochastic Parity Games
复制标题

DOI:
10.4230/lipics.concur.2020.21
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
E. M. Hahn;Mateo Perez;S. Schewe;F. Somenzi;Ashutosh Trivedi;D. Wojtczak
E. M. Hahn;Mateo Perez;S. Schewe;F. Somenzi;Ashutosh Trivedi;D. Wojtczak
中科院分区:
其他
文献类型:
--
作者:
E. M. Hahn;Mateo Perez;S. Schewe;F. Somenzi;Ashutosh Trivedi;D. Wojtczak

文献摘要

相似文献

本文研究了使用无模型强化学习来计算具有同等目标的两人随机博弈的最优值。在此设置中,两个决策者(玩家 Min 和玩家 Max)在有限博弈竞技场(具有未知但固定概率分布的随机博弈图)上竞争,以分别最小化和最大化满足平价目标的概率。我们将随机奇偶博弈简化为带有参数 ε 的一系列随机可达博弈,使得当参数 ε 趋向于 0 时,随机奇偶博弈的值等于相应简单随机博弈的值的极限。由于这种简化不需要了解底层博弈领域的概率转移结构,因此可以使用无模型强化学习算法(例如 minimax Q-learning)来逼近该值和相互的最佳响应潜在随机平价博弈中双方玩家的策略。我们还提出了从 1 12 名玩家平价游戏到可达性游戏的简化简化,避免诉诸非确定性。最后,我们报告了这两种减少的实验评估。
This paper investigates the use of model-free reinforcement learning to compute the optimal value in two-player stochastic games with parity objectives. In this setting, two decision makers, player Min and player Max, compete on a finite game arena – a stochastic game graph with unknown but fixed probability distributions – to minimize and maximize, respectively, the probability of satisfying a parity objective. We give a reduction from stochastic parity games to a family of stochastic reachability games with a parameter ε , such that the value of a stochastic parity game equals the limit of the values of the corresponding simple stochastic games as the parameter ε tends to 0. Since this reduction does not require the knowledge of the probabilistic transition structure of the underlying game arena, model-free reinforcement learning algorithms, such as minimax Q-learning, can be used to approximate the value and mutual best-response strategies for both players in the underlying stochastic parity game. We also present a streamlined reduction from 1 12 -player parity games to reachability games that avoids recourse to nondeterminism. Finally, we report on the experimental evaluations of both reductions.