Influencing Towards Stable Multi-Agent Interactions

Influencing Towards Stable Multi-Agent Interactions
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Woodrow Z. Wang;Andy Shih;Annie Xie;Dorsa Sadigh
Woodrow Z. Wang;Andy Shih;Annie Xie;Dorsa Sadigh
中科院分区:
其他
文献类型:
--
作者:
Woodrow Z. Wang;Andy Shih;Annie Xie;Dorsa Sadigh

文献摘要

相似文献

在多智能体环境中学习是困难的,由于对手或合作伙伴的行为变化所引入的非平稳性。而不是被动地适应其他代理(对手或合作伙伴)的行为,我们提出了一种算法,主动影响其他代理的策略,以稳定-这可以抑制非平稳性所造成的其他代理。我们学习其他代理的策略的低维潜在表示以及潜在策略如何随我们的机器人行为而演变的动态。有了这个学习的动态模型,我们可以定义一个无监督的稳定性奖励,来训练我们的机器人故意影响其他智能体,使其朝着单一策略稳定下来。我们证明了稳定的有效性,在各种模拟环境中,包括自动驾驶,紧急通信和机器人操作,以提高效率,最大限度地提高任务奖励。我们在我们的网站上展示定性结果:https://sites.google.com/view/stable-marl/。
Learning in multi-agent environments is difficult due to the non-stationarity introduced by an opponent's or partner's changing behaviors. Instead of reactively adapting to the other agent's (opponent or partner) behavior, we propose an algorithm to proactively influence the other agent's strategy to stabilize -- which can restrain the non-stationarity caused by the other agent. We learn a low-dimensional latent representation of the other agent's strategy and the dynamics of how the latent strategy evolves with respect to our robot's behavior. With this learned dynamics model, we can define an unsupervised stability reward to train our robot to deliberately influence the other agent to stabilize towards a single strategy. We demonstrate the effectiveness of stabilizing in improving efficiency of maximizing the task reward in a variety of simulated environments, including autonomous driving, emergent communication, and robotic manipulation. We show qualitative results on our website: https://sites.google.com/view/stable-marl/.