Hybrid Independent Learning in Cooperative Markov Games
Hybrid Independent Learning in Cooperative Markov Games
复制标题
DOI:
10.1007/978-3-030-64096-5_6
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Roi Yehoshua;Chris Amato
中科院分区:
文献类型:
--
作者:
Roi Yehoshua;Chris Amato
Independent agents learning by reinforcement must overcome several difficulties, including non-stationarity, miscoordination, and relative overgeneralization. An independent learner may receive different rewards for the same state and action at different time steps, depending on the actions of the other agents in that state. Existing multi-agent learning methods try to overcome these issues by using various techniques, such as hysteresis or leniency. However, they all use the latest reward signal to update the Q function. Instead, we propose to keep track of the rewards received for each state-action pair, and use a hybrid approach for updating the Q values: the agents initially adopt an optimistic disposition by using the maximum reward observed, and then transform into average reward learners. We show both analytically and empirically that this technique can improve the convergence and stability of the learning, and is able to deal robustly with overgeneralization, miscoordination, and high degree of stochasticity in the reward and transition functions. Our method outperforms state-of-the-art multi-agent learning algorithms across a spectrum of stochastic and partially observable games, while requiring little parameter tuning.