Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games
Off-Policy Exploitability-Evaluation in Two-Player Zero-Sum Markov Games
复制标题
二人零和马尔可夫博弈中的离策略可利用性评估
DOI:
10.5555/3463952.3463968
复制
发表时间:
2021
影响因子:
6
通讯作者:
Yusuke Kaneko
中科院分区:
文献类型:
--
作者:
Kenshi Abe;Yusuke Kaneko
Off-policy evaluation (OPE) is the problem of evaluating new policies using historical data obtained from a different policy. In the recent OPE context, most studies have focused on single-player cases, and not on multi-player cases. In this study, we propose OPE estimators constructed by the doubly robust and double reinforcement learning estimators in two-player zero-sum Markov games. The proposed estimators project exploitability that is often used as a metric for determining how close a policy profile (i.e., a tuple of policies) is to a Nash equilibrium in two-player zero-sum games. We prove the exploitability estimation error bounds for the proposed estimators. We then propose the methods to find the best candidate policy profile by selecting the policy profile that minimizes the estimated exploitability from a given policy profile class. We prove the regret bounds of the policy profiles selected by our methods. Finally, we demonstrate the effectiveness and performance of the proposed estimators through experiments.
DOI:
--
发表时间:
2019
期刊:
Advances in neural information processing systems
影响因子:
--
作者:
Kallus, Nathan;Uehara, Masatoshi
通讯作者:
Uehara, Masatoshi
DOI:
--
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
作者:
Carlos Martin;T. Sandholm
通讯作者:
Carlos Martin;T. Sandholm
DOI:
--
发表时间:
2020-02
期刊:
--
影响因子:
--
作者:
Nathan Kallus;Masatoshi Uehara
通讯作者:
Nathan Kallus;Masatoshi Uehara