Efficient Deviation Types and Learning for Hindsight Rationality in Extensive-Form Games

Efficient Deviation Types and Learning for Hindsight Rationality in Extensive-Form Games
复制标题

DOI:
10.48550/arxiv.2205.12031
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Dustin Morrill;Ryan D'Orazio;Marc Lanctot;J. R. Wright;Michael H. Bowling;A. Greenwald
Dustin Morrill;Ryan D'Orazio;Marc Lanctot;J. R. Wright;Michael H. Bowling;A. Greenwald
中科院分区:
其他
文献类型:
--
作者:
Dustin Morrill;Ryan D'Orazio;Marc Lanctot;J. R. Wright;Michael H. Bowling;A. Greenwald

文献摘要

相似文献

后见之明理性(英语:Hindsight rationality)是一种玩一般和游戏的方法,它规定了个体代理关于一组偏差的无遗憾学习动态,并进一步描述了多个代理之间的联合理性行为。为了在顺序决策环境中发展后见之明的理性学习,我们将行为偏差形式化为尊重扩展形式游戏结构的一般偏差类。将时间选择的思想集成到反事实遗憾最小化(CFR)中,我们引入了扩展形式的遗憾最小化(EFR)算法,该算法可以对任何给定的行为偏差集实现事后理性,其计算与集合的复杂性密切相关。我们识别出行为偏差子集,即部分序列偏差类型,这些偏差类型继承了先前研究的类型,并在中等长度的游戏中导致有效的EFR实例。此外,我们对基准游戏中不同偏离类型的EFR实例进行了全面的实证分析,我们发现更强的类型通常会带来更好的性能。
Hindsight rationality is an approach to playing general-sum games that prescribes no-regret learning dynamics for individual agents with respect to a set of deviations, and further describes jointly rational behavior among multiple agents with mediated equilibria. To develop hindsight rational learning in sequential decision-making settings, we formalize behavioral deviations as a general class of deviations that respect the structure of extensive-form games. Integrating the idea of time selection into counterfactual regret minimization (CFR), we introduce the extensive-form regret minimization (EFR) algorithm that achieves hindsight rationality for any given set of behavioral deviations with computation that scales closely with the complexity of the set. We identify behavioral deviation subsets, the partial sequence deviation types, that subsume previously studied types and lead to efficient EFR instances in games with moderate lengths. In addition, we present a thorough empirical analysis of EFR instantiated with different deviation types in benchmark games, where we find that stronger types typically induce better performance.