Can bounded and self-interested agents be teammates? Application to planning in ad hoc teams

Can bounded and self-interested agents be teammates? Application to planning in ad hoc teams
复制标题

有界和自利的代理人可以成为队友吗?

DOI:
10.1007/s10458-016-9354-4
复制
发表时间:
2016-11
影响因子:
1.9
通讯作者:
Yingke Chen
Yingke Chen
中科院分区:
计算机科学4区
文献类型:
--
作者:
Muthukumaran Ch;rasekaran;Prashant Doshi;Yifeng Zeng;Yingke Chen

文献摘要

参考文献

相似文献

特别团队的规划是具有挑战性的,因为它涉及代理在没有任何事先协调或沟通的情况下进行协作。重点是原则性的方法,为一个单一的代理人与他人合作。这激发了在自利决策框架的背景下调查特设团队合作问题。在多智能体环境中,参与个体决策的智能体面临着必须对其他智能体的行为进行推理的任务,这反过来又可能涉及对其他智能体的推理。一个已建立的近似操作这种方法是从下面的无限嵌套引入0级模型。为了这项研究的目的,在多智能体设置的个人,自利的决策建模使用交互式动态影响图(I-DID)。这些是图形模型,其好处是它们自然地提供了问题的因子表示,允许代理将动态模型归因于其他人并对其进行推理。我们证明了一个有界的,有限嵌套的推理由一个自私的代理的含义是,我们可能不会获得最佳的团队解决方案,在合作的设置,如果它是一个团队的一部分。我们通过包括0级模型来解决这个限制,这些模型的解决方案涉及强化学习。我们展示了如何学习集成到规划的背景下,I-DID。这有利于最佳的队友行为,我们证明了它的适用性,特设团队合作的几个问题域和配置。
Planning for ad hoc teamwork is challenging because it involves agents collaborating without any prior coordination or communication. The focus is on principled methods for a single agent to cooperate with others. This motivates investigating the ad hoc teamwork problem in the context of self-interested decision-making frameworks. Agents engaged in individual decision making in multiagent settings face the task of having to reason about other agents’ actions, which may in turn involve reasoning about others. An established approximation that operationalizes this approach is to bound the infinite nesting from below by introducing level 0 models. For the purposes of this study, individual, self-interested decision making in multiagent settings is modeled using interactive dynamic influence diagrams (I-DID). These are graphical models with the benefit that they naturally offer a factored representation of the problem, allowing agents to ascribe dynamic models to others and reason about them. We demonstrate that an implication of bounded, finitely-nested reasoning by a self-interested agent is that we may not obtain optimal team solutions in cooperative settings, if it is part of a team. We address this limitation by including models at level 0 whose solutions involve reinforcement learning. We show how the learning is integrated into planning in the context of I-DIDs. This facilitates optimal teammate behavior, and we demonstrate its applicability to ad hoc teamwork on several problem domains and configurations.
DOI: 10.1145/1558109.1558138
发表时间: 2009-05
期刊: Epigenetics
影响因子: 3.7
作者:
Prashant Doshi;Yi-feng Zeng
通讯作者: Prashant Doshi;Yi-feng Zeng
DOI: 10.1145/2600057.2602907
发表时间: 2014-06
期刊: Proceedings of the fifteenth ACM conference on Economics and computation
影响因子: --
作者:
J. R. Wright;Kevin Leyton-Brown
通讯作者: J. R. Wright;Kevin Leyton-Brown
DOI: 10.1007/978-3-642-15117-0_10
发表时间: 2009-05
期刊: --
影响因子: --
作者:
P. Stone;G. Kaminka;J. Rosenschein
通讯作者: P. Stone;G. Kaminka;J. Rosenschein
DOI: 10.1109/wi-iat.2010.74
发表时间: 2010-08
期刊: 2010 IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology
影响因子: --
作者:
Prashant Doshi;Muthukumaran Chandrasekaran;Yi-feng Zeng
通讯作者: Prashant Doshi;Muthukumaran Chandrasekaran;Yi-feng Zeng