Inducing Social Optimality in Games via Adaptive Incentive Design

Inducing Social Optimality in Games via Adaptive Incentive Design
复制标题

DOI:
10.1109/cdc51059.2022.9992685
复制
发表时间:
2022-04
期刊:
2022 IEEE 61st Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
C. Maheshwari;Kshitij Kulkarni;Manxi Wu;S. Sastry
C. Maheshwari;Kshitij Kulkarni;Manxi Wu;S. Sastry
中科院分区:
其他
文献类型:
--
作者:
C. Maheshwari;Kshitij Kulkarni;Manxi Wu;S. Sastry

文献摘要

被引文献

相似文献

社会规划者如何适应性地激励在战略环境中学习的自私主体,以实现长期的社会最优结果?我们提出了两个时间尺度的学习动态来回答游戏中的这个问题。在我们的学习动态中,玩家采用一类学习规则以更快的时间尺度更新他们的策略,而社会规划者以更慢的时间尺度更新激励机制。具体来说,激励机制的更新是基于每个参与者的外部性,其评估为每个时间步内参与者的边际成本与社会边际成本之间的差值。我们表明,我们的学习动态的任何固定点都对应于最优激励机制,使得相应的纳什均衡也实现社会最优。我们还为学习动态收敛到固定点提供充分的条件,以便自适应激励机制最终产生社会最优结果。最后,作为一个例子,我们证明了在有限参与者的古诺竞争中满足收敛的充分条件。
How can a social planner adaptively incentivize selfish agents who are learning in a strategic environment to induce a socially optimal outcome in the long run? We propose a two-timescale learning dynamics to answer this question in games. In our learning dynamics, players adopt a class of learning rules to update their strategies at a faster timescale, while a social planner updates the incentive mechanism at a slower timescale. In particular, the update of the incentive mechanism is based on each player’s externality, which is evaluated as the difference between the player’s marginal cost and the society’s marginal cost in each time step. We show that any fixed point of our learning dynamics corresponds to the optimal incentive mechanism such that the corresponding Nash equilibrium also achieves social optimality. We also provide sufficient conditions for the learning dynamics to converge to a fixed point so that the adaptive incentive mechanism eventually induces a socially optimal outcome. Finally, as an example, we demonstrate that the sufficient conditions for convergence are satisfied in Cournot competition with finite players.