Hybrid Independent Learning in Cooperative Markov Games

Hybrid Independent Learning in Cooperative Markov Games
复制标题

DOI:
10.1007/978-3-030-64096-5_6
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Roi Yehoshua;Chris Amato
Roi Yehoshua;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Roi Yehoshua;Chris Amato

文献摘要

相似文献

通过强化学习的独立代理必须克服几个困难,包括非平稳性,不协调和相对过度概括。独立学习者可能会在不同的时间步长对相同的状态和动作获得不同的奖励,这取决于该状态下其他智能体的动作。现有的多智能体学习方法试图克服这些问题,通过使用各种技术,如滞后或宽大。然而,它们都使用最新的奖励信号来更新Q函数。相反,我们建议跟踪每个状态-动作对收到的奖励,并使用混合方法更新Q值:代理最初通过使用观察到的最大奖励采取乐观的处置,然后转化为平均奖励学习者。我们的分析和经验表明,这种技术可以提高学习的收敛性和稳定性,并能够稳健地处理过度泛化,不协调和高度的随机性的奖励和过渡函数。我们的方法优于国家的最先进的多智能体学习算法在频谱的随机和部分可观察的游戏,而需要很少的参数调整。
Independent agents learning by reinforcement must overcome several difficulties, including non-stationarity, miscoordination, and relative overgeneralization. An independent learner may receive different rewards for the same state and action at different time steps, depending on the actions of the other agents in that state. Existing multi-agent learning methods try to overcome these issues by using various techniques, such as hysteresis or leniency. However, they all use the latest reward signal to update the Q function. Instead, we propose to keep track of the rewards received for each state-action pair, and use a hybrid approach for updating the Q values: the agents initially adopt an optimistic disposition by using the maximum reward observed, and then transform into average reward learners. We show both analytically and empirically that this technique can improve the convergence and stability of the learning, and is able to deal robustly with overgeneralization, miscoordination, and high degree of stochasticity in the reward and transition functions. Our method outperforms state-of-the-art multi-agent learning algorithms across a spectrum of stochastic and partially observable games, while requiring little parameter tuning.