Commitment-driven distributed joint policy search

Commitment-driven distributed joint policy search
复制标题

承诺驱动的分布式联合政策搜索

DOI:
10.1145/1329125.1329216
复制
发表时间:
2007
影响因子:
1.9
通讯作者:
E. Durfee
E. Durfee
中科院分区:
计算机科学4区
文献类型:
--
作者:
S. Witwicki;E. Durfee

文献摘要

被引文献

相似文献

去中心化的MDP提供了多智能体环境中强大的交互模型,但通常很难甚至在计算上无法最佳解决。在这里,我们开发了一个层次化的方法来解决一组有限的分散MDPs。通过与其他代理人形成承诺,并在其本地MDP中简明地建模这些承诺,代理人有效地,高效地,分布式地制定协调的本地政策。我们引入了一种新的建设,捕捉承诺的约束,当地的政策,并显示如何线性规划可以用来实现局部最优受这些约束。与其他承诺执行方法相比,我们表明我们在捕获预期的承诺语义,同时最大限度地提高局部效用方面更强大。我们还描述了一个承诺空间的启发式搜索算法,可以用来近似最优的联合政策。初步的经验评估表明,我们的方法产生更快的近似解比传统的编码的问题,作为一个多智能体MDP将允许,当包裹在一个详尽的承诺空间搜索,将找到最佳的全局解决方案。
Decentralized MDPs provide powerful models of interactions in multiagent environments, but are often very difficult or even computationally infeasible to solve optimally. Here we develop a hierarchical approach to solving a restricted set of decentralized MDPs. By forming commitments with other agents and modeling these concisely in their local MDPs, agents effectively, efficiently, and distributively formulate co-ordinated local policies. We introduce a novel construction that captures commitments as constraints on local policies and show how Linear Programming can be used to achieve local optimality subject to these constraints. In contrast to other commitment enforcement approaches, we show ours to be more robust in capturing the intended commitment semantics while maximizing local utility. We also describe a commitment-space heuristic search algorithm that can be used to approximate optimal joint policies. A preliminary empirical evaluation suggests that our approach yields faster approximate solutions than the conventional encoding of the problem as a multiagent MDP would allow and, when wrapped in an exhaustive commitment-space search, will find the optimal global solution.