Conditional Imitation Learning for Multi-Agent Games

Conditional Imitation Learning for Multi-Agent Games
复制标题

DOI:
10.1109/hri53351.2022.9889671
复制
发表时间:
2022-01
期刊:
2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI)
影响因子:
--
通讯作者:
Andy Shih;Stefano Ermon;Dorsa Sadigh
Andy Shih;Stefano Ermon;Dorsa Sadigh
中科院分区:
其他
文献类型:
--
作者:
Andy Shih;Stefano Ermon;Dorsa Sadigh

文献摘要

相似文献

虽然多智能体学习的进步已经使得越来越复杂的代理的培训成为可能,但大多数现有技术产生的最终政策并不是为了适应新合作伙伴的策略而设计的。然而,我们希望我们的人工智能代理能够根据周围人的策略来调整自己的策略。在这项工作中,我们研究了有条件的多智能体模仿学习的问题,我们在训练时可以获得联合轨迹演示,并且我们必须在测试时与新的合作伙伴进行交互和适应。这种设置是具有挑战性的,因为我们必须推断新合作伙伴的策略,并根据该策略调整我们的策略,所有这些都不需要了解环境奖励或动态。我们正式提出了这个问题的条件多智能体模仿学习,并提出了一种新的方法来解决的可扩展性和数据稀缺的困难。我们的关键见解是,在多智能体游戏的合作伙伴之间的变化往往是高度结构化的,可以通过一个低秩子空间表示。利用张量分解的工具,我们的模型在自我和合作伙伴代理策略上学习一个低秩子空间,然后通过在子空间中插值来推断和适应新的合作伙伴策略。我们实验了各种协作任务,包括强盗、粒子和Hanabi环境。此外,我们测试我们的条件政策对真实的人类合作伙伴在用户研究的过度烹饪游戏。与基线相比,我们的模型更好地适应新的合作伙伴,并鲁棒地处理各种设置,从离散/连续动作到人工智能/人类合作伙伴的静态/在线评估。
While advances in multi-agent learning have enabled the training of increasingly complex agents, most existing techniques produce a final policy that is not designed to adapt to a new partner's strategy. However, we would like our AI agents to adjust their strategy based on the strategies of those around them. In this work, we study the problem of conditional multi-agent imitation learning, where we have access to joint trajectory demonstrations at training time, and we must interact with and adapt to new partners at test time. This setting is challenging because we must infer a new partner's strategy and adapt our policy to that strategy, all without knowledge of the environment reward or dynamics. We formalize this problem of conditional multi-agent imitation learning, and propose a novel approach to address the difficulties of scalability and data scarcity. Our key insight is that variations across partners in multi-agent games are often highly structured, and can be represented via a low-rank subspace. Leveraging tools from tensor decomposition, our model learns a low-rank subspace over ego and partner agent strategies, then infers and adapts to a new partner strategy by interpolating in the subspace. We experiments with a mix of collaborative tasks, including bandits, particle, and Hanabi environments. Additionally, we test our conditional policies against real human partners in a user study on the Overcooked game. Our model adapts better to new partners compared to baselines, and robustly handles diverse settings ranging from discrete/continuous actions and static/online evaluation with AI/human partners.