Multiagent Evaluation under Incomplete Information

Multiagent Evaluation under Incomplete Information
复制标题

不完全信息下的多主体评估

DOI:
10.1145/3477045
复制
发表时间:
2019
期刊:
ArXiv
影响因子:
--
通讯作者:
R. Munos
R. Munos
中科院分区:
--
文献类型:
--
作者:
Mark Rowland;Shayegan Omidshafiei;K. Tuyls;J. Pérolat;Michal Valko;G. Piliouras;R. Munos

文献摘要

参考文献

被引文献

相似文献

本文研究了在不完全信息环境下多智能体学习策略的评价问题,它在智能体的排序和训练中起着关键作用。传统上,研究人员依赖于Elo评级来实现这一目的,最近的研究也使用了基于纳什均衡的方法。不幸的是,Elo无法处理不可传递的代理交互,其他技术仅限于零和,两个玩家设置或受到纳什均衡难以计算的事实的限制。最近,排名方法称为$\alpha$-排名,依赖于一个新的基于图形的博弈论解决方案的概念,被证明可追溯地适用于一般的游戏。然而,基于Elo或$\alpha$-Rank的评估通常假设无噪声的游戏结果,尽管数据通常是从有噪声的模拟中收集的,这使得这种假设在实践中不切实际。本文研究了多智能体评价在不完全信息制度,涉及一般和多玩家游戏与噪声的结果。我们得到样本的复杂性保证,在这种情况下,自信地排名代理。我们提出了自适应算法的准确排名,提供正确性和样本的复杂性保证,然后介绍了一种手段连接的不确定性在嘈杂的匹配结果的不确定性排名。我们评估这些方法在几个领域的性能,包括伯努利游戏,足球元游戏,库恩扑克。
This paper investigates the evaluation of learned multiagent strategies in the incomplete information setting, which plays a critical role in ranking and training of agents. Traditionally, researchers have relied on Elo ratings for this purpose, with recent works also using methods based on Nash equilibria. Unfortunately, Elo is unable to handle intransitive agent interactions, and other techniques are restricted to zero-sum, two-player settings or are limited by the fact that the Nash equilibrium is intractable to compute. Recently, a ranking method called $\alpha$-Rank, relying on a new graph-based game-theoretic solution concept, was shown to tractably apply to general games. However, evaluations based on Elo or $\alpha$-Rank typically assume noise-free game outcomes, despite the data often being collected from noisy simulations, making this assumption unrealistic in practice. This paper investigates multiagent evaluation in the incomplete information regime, involving general-sum many-player games with noisy outcomes. We derive sample complexity guarantees required to confidently rank agents in this setting. We propose adaptive algorithms for accurate ranking, provide correctness and sample complexity guarantees, then introduce a means of connecting uncertainties in noisy match outcomes to uncertainties in rankings. We evaluate the performance of these approaches in several domains, including Bernoulli games, a soccer meta-game, and Kuhn poker.
DOI: 10.1145/2482540.2482558
发表时间: 2013-02
期刊: J. Mach. Learn. Res.
影响因子: --
作者:
John Fearnley;Martin Gairing;P. Goldberg;Rahul Savani
通讯作者: John Fearnley;Martin Gairing;P. Goldberg;Rahul Savani