Learning in Stackelberg Games with Non-myopic Agents

Learning in Stackelberg Games with Non-myopic Agents
复制标题

与非近视智能体一起在 Stackelberg 游戏中学习

DOI:
10.1145/3490486.3538308
复制
发表时间:
2022
期刊:
Proceedings of the 23rd ACM Conference on Economics and Computation
影响因子:
--
通讯作者:
Wei, Alexander
Wei, Alexander
中科院分区:
--
文献类型:
--
作者:
Haghtalab, Nika;Lykouris, Thodoris;Nietert, Sloan;Wei, Alexander

文献摘要

参考文献

被引文献

相似文献

Stackelberg博弈是战略委托-代理互动的典型模型。例如,考虑一个防御系统,它在攻击执行之前将其安全资源分配到高风险目标;或者一个税务政策制定者在看到提交的税务报告之前制定了何时触发审计的规则;或者一个卖家在知道客户的购买倾向之前选择了一个价格。在这些场景中的每一个中,委托人首先选择操作x∈X,然后代理使用操作y∈Y进行反应,其中X和Y分别是委托人和代理的操作空间。在上面的示例中,代理操作分别对应于要攻击的目标、为逃避审计而支付的税款和购买的金额。通常,委托人希望x在代理人做出最佳反应y=br(X)时最大化他们的收益;这样的对(x,y)是Stackelberg均衡。通过承诺一种策略,委托人可以保证他们获得比相应的同时博弈的固定点均衡更高的回报。然而,找到这样的策略需要知道代理的收益函数,当面对未知的代理收益时,委托人可以尝试通过与代理的反复交互来学习最佳响应。如果(天真的)代理人没有意识到这样的学习发生,并且总是发挥最好的反应,委托人可以使用经典的在线学习方法来优化他们自己在阶段游戏中的收益。从近视代理学习已经在多个Stackelberg游戏中进行了广泛的研究,包括安全游戏[2,6,7],按需学习[1,5]和策略分类[3,4]。然而,长期生存的代理通常不会自愿提供未来可能对他们不利的信息。在在线环境中尤其如此,在这种环境中,学习者试图尽快利用最近学习到的行为模式,代理可以看到偏离其瞬时最佳响应并将学习者引向歧途的有形优势。这种在学习算法的(统计)效率和它们可能在长期内创造的反常激励之间的权衡,将我们带入了这项工作的主要问题:在一般的Stackelberg游戏中,针对非近视代理人的学习有哪些原则性的方法?如何将从对抗近视剂学习中获得的见解应用到非近视病例的学习中?
Stackelberg games are a canonical model for strategic principal-agent interactions. Consider, for instance, a defense system that distributes its security resources across high-risk targets prior to attacks being executed; or a tax policymaker who sets rules on when audits are triggered prior to seeing filed tax reports; or a seller who chooses a price prior to knowing a customer's proclivity to buy. In each of these scenarios, a principal first selects an action x∈X and then an agent reacts with an action y∈Y, where X and Y are the principal's and agent's action spaces, respectively. In the examples above, agent actions correspond to which target to attack, how much tax to pay to evade an audit, and how much to purchase, respectively. Typically, the principal wants an x that maximizes their payoff when the agent plays a best response y = br(x); such a pair (x, y) is a Stackelberg equilibrium. By committing to a strategy, the principal can guarantee they achieve a higher payoff than in the fixed point equilibrium of the corresponding simultaneous-play game. However, finding such a strategy requires knowledge of the agent's payoff function.When faced with unknown agent payoffs, the principal can attempt to learn a best response via repeated interactions with the agent. If a (naïve) agent is unaware that such learning occurs and always plays a best response, the principal can use classical online learning approaches to optimize their own payoff in the stage game. Learning from myopic agents has been extensively studied in multiple Stackelberg games, including security games[2,6,7], demand learning[1,5], and strategic classification[3,4].However, long-lived agents will generally not volunteer information that can be used against them in the future. This is especially the case in online environments where a learner seeks to exploit recently learned patterns of behavior as soon as possible, and the agent can see a tangible advantage for deviating from its instantaneous best response and leading the learner astray. This trade-off between the (statistical) efficiency of learning algorithms and the perverse incentives they may create over the long-term brings us to the main questions of this work: What are principled approaches to learning against non-myopic agents in general Stackelberg games? How can insights from learning against myopic agents be applied to learning in the non-myopic case?
DOI: --
发表时间: 2017-09
期刊: --
影响因子: --
作者:
Ciara Pike-Burke;Shipra Agrawal;Csaba Szepesvari;S. Grünewälder
通讯作者: Ciara Pike-Burke;Shipra Agrawal;Csaba Szepesvari;S. Grünewälder
DOI: --
发表时间: 2006-12
期刊: J. Mach. Learn. Res.
影响因子: --
作者:
Eyal Even-Dar;Shie Mannor;Y. Mansour
通讯作者: Eyal Even-Dar;Shie Mannor;Y. Mansour
学习针对非短视投标人的最优保留价
DOI: --
发表时间: 2018
期刊: Neural Information Processing Systems
影响因子: --
作者:
Zhiyi Huang;Jinyan Liu;Xiangning Wang
通讯作者: Xiangning Wang
动态激励感知学习:关联拍卖中的稳健定价
DOI: 10.1287/opre.2020.1991
发表时间: 2021
影响因子: 2.7
作者:
Negin Golrezaei, Adel Javanmard
通讯作者: Negin Golrezaei, Adel Javanmard
DOI: 10.1287/opre.2020.2007
发表时间: 2021
影响因子: 2.7
作者:
Kanoria, Yash;Nazerzadeh, Hamid
通讯作者: Nazerzadeh, Hamid