Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games

Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games
复制标题

二人零和马尔可夫博弈中的离策略可利用性评估

DOI:
10.5555/3463952.3463968
复制
发表时间:
2021
影响因子:
6
通讯作者:
Yusuke Kaneko
Yusuke Kaneko
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kenshi Abe;Yusuke Kaneko

文献摘要

参考文献

被引文献

相似文献

政策外评估(OPE)是使用从不同政策获得的历史数据来评估新政策的问题。在最近的OPE背景下,大多数研究都集中在单人情况下,而不是多人情况下。在这项研究中,我们提出了由两人零和马尔可夫博弈中的双重鲁棒和双重强化学习估计器构造的OPE估计器。所提出的估计器预测了可利用性,该可利用性通常用作确定策略配置文件(即,一组策略)是两个玩家零和游戏中的纳什均衡。我们证明了所提出的估计的可利用性估计误差界。然后,我们提出的方法来找到最佳的候选人的政策配置文件,选择的政策配置文件,最大限度地减少估计利用从一个给定的政策配置文件类。我们证明了我们的方法选择的政策配置文件的遗憾界。最后,我们通过实验证明了所提出的估计器的有效性和性能。
Off-policy evaluation (OPE) is the problem of evaluating new policies using historical data obtained from a different policy. In the recent OPE context, most studies have focused on single-player cases, and not on multi-player cases. In this study, we propose OPE estimators constructed by the doubly robust and double reinforcement learning estimators in two-player zero-sum Markov games. The proposed estimators project exploitability that is often used as a metric for determining how close a policy profile (i.e., a tuple of policies) is to a Nash equilibrium in two-player zero-sum games. We prove the exploitability estimation error bounds for the proposed estimators. We then propose the methods to find the best candidate policy profile by selecting the policy profile that minimizes the estimated exploitability from a given policy profile class. We prove the regret bounds of the policy profiles selected by our methods. Finally, we demonstrate the effectiveness and performance of the proposed estimators through experiments.
本质上高效、稳定且有界的强化学习离策略评估
DOI: --
发表时间: 2019
期刊: Advances in neural information processing systems
影响因子: --
作者:
Kallus, Nathan;Uehara, Masatoshi
通讯作者: Uehara, Masatoshi
DOI: --
发表时间: 2020-02
期刊: ArXiv
影响因子: --
作者:
Carlos Martin;T. Sandholm
通讯作者: Carlos Martin;T. Sandholm
DOI: --
发表时间: 2020-02
期刊: --
影响因子: --
作者:
Nathan Kallus;Masatoshi Uehara
通讯作者: Nathan Kallus;Masatoshi Uehara