Deep Q-Learning for Nash Equilibria: Nash-DQN

Deep Q-Learning for Nash Equilibria: Nash-DQN
复制标题

纳什均衡的深度 Q 学习:Nash-DQN

DOI:
10.1080/1350486x.2022.2136727
复制
发表时间:
2019
影响因子:
--
通讯作者:
S. Jaimungal
S. Jaimungal
中科院分区:
--
文献类型:
--
作者:
P. Casgrain;Brian Ning;S. Jaimungal

文献摘要

被引文献

相似文献

多智能体随机博弈的无模型学习是一个活跃的研究领域。然而,现有的强化学习算法通常仅限于零和游戏,并且仅适用于小的状态动作空间或其他简化设置。在这里,我们开发了一种新的数据高效的Deep-Q-learning方法,用于一般和随机游戏的纳什均衡的无模型学习。该算法使用了局部线性二次扩展的随机游戏,从而导致解析可解的最佳行动。扩展由深度神经网络参数化,以提供足够的灵活性来学习环境,而无需体验所有状态-动作对。我们研究的对称性的算法源于标签不变的随机游戏,并作为一个概念的证明,我们的算法学习在竞争激烈的电子市场中的最佳交易策略。
ABSTRACT Model-free learning for multi-agent stochastic games is an active area of research. Existing reinforcement learning algorithms, however, are often restricted to zero-sum games and are applicable only in small state-action spaces or other simplified settings. Here, we develop a new data-efficient Deep-Q-learning methodology for model-free learning of Nash equilibria for general-sum stochastic games. The algorithm uses a locally linear-quadratic expansion of the stochastic game, which leads to analytically solvable optimal actions. The expansion is parametrized by deep neural networks to give it sufficient flexibility to learn the environment without the need to experience all state-action pairs. We study symmetry properties of the algorithm stemming from label-invariant stochastic games and as a proof of concept, apply our algorithm to learning optimal trading strategies in competitive electronic markets.