Difference rewards policy gradients

Difference rewards policy gradients
复制标题

差异奖励政策梯度

DOI:
10.1007/s00521-022-07960-5
复制
发表时间:
2022
影响因子:
6
通讯作者:
Castellini J
Castellini J
中科院分区:
计算机科学3区
文献类型:
--
作者:
Castellini J

文献摘要

相似文献

策略梯度方法已经成为多智能体强化学习中最流行的一类算法。然而,一个关键的挑战,这是没有解决的许多这些方法是多代理信用分配:评估代理的整体性能,这是至关重要的学习良好的政策的贡献。我们提出了一种名为Dr.Reinforce的新算法,该算法通过将差异奖励与策略梯度相结合来明确解决这个问题,以便在奖励函数已知时学习分散策略。通过直接区分奖励函数,Reinforce博士避免了与学习Q函数相关的困难,这是由反事实多代理策略梯度(COMA)完成的,COMA是一种最先进的差异奖励方法。对于奖励函数未知的应用程序,我们展示了Reinforce博士的一个版本的有效性,该版本学习了一个用于估计差异奖励的额外奖励网络。
Policy gradient methods have become one of the most popular classes of algorithms for multi-agent reinforcement learning. A key challenge, however, that is not addressed by many of these methods is multi-agent credit assignment: assessing an agent’s contribution to the overall performance, which is crucial for learning good policies. We propose a novel algorithm called Dr.Reinforce that explicitly tackles this by combining difference rewards with policy gradients to allow for learning decentralized policies when the reward function is known. By differencing the reward function directly, Dr.Reinforce avoids difficulties associated with learning theQ-function as done by counterfactual multi-agent policy gradients (COMA), a state-of-the-art difference rewards method. For applications where the reward function is unknown, we show the effectiveness of a version of Dr.Reinforce that learns an additional reward network that is used to estimate the difference rewards.