Dirichlet-Multinomial Counterfactual Rewards for Heterogeneous Multiagent Systems

Dirichlet-Multinomial Counterfactual Rewards for Heterogeneous Multiagent Systems
复制标题

DOI:
10.1109/mrs.2019.8901077
复制
发表时间:
2019-08
期刊:
2019 International Symposium on Multi-Robot and Multi-Agent Systems (MRS)
影响因子:
--
通讯作者:
Gaurav Dixit;Nicholas Zerbel;Kagan Tumer
Gaurav Dixit;Nicholas Zerbel;Kagan Tumer
中科院分区:
其他
文献类型:
--
作者:
Gaurav Dixit;Nicholas Zerbel;Kagan Tumer

文献摘要

相似文献

多机器人团队已被证明可以有效地完成需要团队成员之间紧密协调的复杂任务。在同质系统中,最近的工作表明,即使目标的代理间耦合要求未得到满足,“踏脚石”奖励也是一种向代理提供潜在有价值的行为反馈的有效方法。在这项工作中,我们提出了一种新机制,用于推断紧密耦合的异构系统中的假设伙伴,称为狄利克雷多项反事实选择(DMCS)。使用 DMCS,我们表明智能体可以学习推断适当的反事实伙伴,通过在修改后的多漫游器探索问题中进行测试来获得更多信息丰富的踏脚石奖励。我们还表明,DMCS 的性能比随机合作伙伴选择基线高出 40% 以上,并且我们还演示了如何使用领域知识来诱导先验,以指导代理学习过程。最后,我们表明,与快速下降的基准性能相比,DMCS 在多达 15 种不同的流动站类型中保持了卓越的性能。
Multi-robot teams have been shown to be effective in accomplishing complex tasks which require tight coordination among team members. In homogeneous systems, recent work has demonstrated that “stepping stone” rewards are an effective way to provide agents with feedback on potentially valuable actions even when the agent-to-agent coupling requirements of an objective are not satisfied. In this work, we propose a new mechanism for inferring hypothetical partners in tightly-coupled, heterogeneous systems called Dirichlet-Multinomial Counterfactual Selection (DMCS). Using DMCS, we show that agents can learn to infer appropriate counterfactual partners to receive more informative stepping stone rewards by testing in a modified multi-rover exploration problem. We also show that DMCS outperforms a random partner selection baseline by over 40%, and we demonstrate how domain knowledge can be used to induce a prior to guide the agent learning process. Finally, we show that DMCS maintains superior performance for up to 15 distinct rover types compared to the performance of the baseline which degrades rapidly.