Leading a Best-Response Teammate in an Ad Hoc Team

Leading a Best-Response Teammate in an Ad Hoc Team
复制标题

DOI:
10.1007/978-3-642-15117-0_10
复制
发表时间:
2009-05
期刊:
--
影响因子:
--
通讯作者:
P. Stone;G. Kaminka;J. Rosenschein
P. Stone;G. Kaminka;J. Rosenschein
中科院分区:
其他
文献类型:
--
作者:
P. Stone;G. Kaminka;J. Rosenschein

文献摘要

被引文献

相似文献

代理团队可能并不总是以有计划的、协调的方式发展。相反,随着部署的代理在电子商务和其他设置中变得越来越常见,以前不熟悉的代理在临时团队设置中合作的机会越来越多。在这样的场景中,在并非所有代理都是完全理性的理念下,单个代理能够与各种可能的队友合作是很有用的。本文考虑了一个要与队友重复交互的代理,该交互将以一种特定的次优但自然的方式适应这种交互。我们用博弈论的术语来形式化这一设置,提供并分析了寻找最优动作序列的完全实现的算法,证明了与这些动作序列的长度有关的一些理论结果,并提供了与我们感兴趣的问题在随机交互环境中的流行度有关的经验结果。
Teams of agents may not always be developed in a planned, coordinated fashion. Rather, as deployed agents become more common in e-commerce and other settings, there are increasing opportunities for previously unacquainted agents to cooperate in ad hoc team settings. In such scenarios, it is useful for individual agents to be able to collaborate with a wide variety of possible teammates under the philosophy that not all agents are fully rational. This paper considers an agent that is to interact repeatedly with a teammate that will adapt to this interaction in a particular suboptimal, but natural way. We formalize this setting in game-theoretic terms, provide and analyze a fully-implemented algorithm for finding optimal action sequences, prove some theoretical results pertaining to the lengths of these action sequences, and provide empirical results pertaining to the prevalence of our problem of interest in random interaction settings.