A framework for sequential planning in multi-agent settings

A framework for sequential planning in multi-agent settings
复制标题

DOI:
10.1613/jair.1579
复制
发表时间:
2005-01-01
影响因子:
5
通讯作者:
Doshi, P
Doshi, P
中科院分区:
计算机科学3区
文献类型:
--
作者:
Gmytrasiewicz, PJ;Doshi, P

文献摘要

被引文献

相似文献

本文将部分可观测马尔可夫决策过程(POMDPs)的框架扩展到多智能体设置,将智能体模型的概念纳入状态空间。代理人保持对环境的物理状态和其他代理人的模型的信念,并且他们使用贝叶斯更新来保持他们的信念。解决方案将信念状态映射到行动。其他代理的模型可能包括他们的信念状态,并与不完全信息游戏中考虑的代理类型有关。我们表示代理人的自主性,假设他们的模型是不直接操纵或观察其他代理人。我们证明了POMDPs的重要性质,如值迭代的收敛性,收敛速度,以及值函数的分段线性和凸性,都可以应用到我们的框架中。我们的方法补充了一个更传统的方法,以互动的设置,使用纳什均衡作为解决方案的范例。我们寻求避免均衡的一些缺点,这些缺点可能是非唯一的并且不能捕获非均衡行为。我们这样做的代价是必须代表,处理和不断修改其他代理商的模型。由于代理人的信念可能是任意嵌套的,决策问题的最优解只能渐近计算。然而,近似信念更新和近似最优计划是可计算的。我们用一个简单的应用领域来说明我们的框架,我们展示了信念更新和价值函数的例子。
This paper extends the framework of partially observable Markov decision processes (POMDPs) to multi-agent settings by incorporating the notion of agent models into the state space. Agents maintain beliefs over physical states of the environment and over models of other agents, and they use Bayesian updates to maintain their beliefs over time. The solutions map belief states to actions. Models of other agents may include their belief states and are related to agent types considered in games of incomplete information. We express the agents' autonomy by postulating that their models are not directly manipulable or observable by other agents. We show that important properties of POMDPs, such as convergence of value iteration, the rate of convergence, and piece-wise linearity and convexity of the value functions carry over to our framework. Our approach complements a more traditional approach to interactive settings which uses Nash equilibria as a solution paradigm. We seek to avoid some of the drawbacks of equilibria which may be non-unique and do not capture off-equilibrium behaviors. We do so at the cost of having to represent, process and continuously revise models of other agents. Since the agent's beliefs may be arbitrarily nested, the optimal solutions to decision making problems are only asymptotically computable. However, approximate belief updates and approximately optimal plans are computable. We illustrate our framework using a simple application domain, and we show examples of belief updates and value functions.