Learning in Zero-Sum Team Markov Games Using Factored Value Functions

Learning in Zero-Sum Team Markov Games Using Factored Value Functions
复制标题

使用分解值函数在零和团队马尔可夫博弈中学习

DOI:
--
复制
发表时间:
2002
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
Ronald E. Parr
Ronald E. Parr
中科院分区:
--
文献类型:
--
作者:
M. Lagoudakis;Ronald E. Parr

文献摘要

被引文献

相似文献

我们提出了一种在零和马尔可夫游戏中学习良好策略的新方法,其中每一方都由多个代理组成,这些代理与对方的代理团队合作。我们的方法在学习过程中需要充分的可观察性和通信性,但学习到的策略可以以分布式方式执行。价值函数表示为因子线性架构,其结构决定了必要的计算资源和通信带宽。这种方法允许在代理之间很少或没有通信的简单表示与代理之间具有广泛协调的复杂、计算密集型表示之间进行权衡。因此,我们提供了一种使用近似来对抗参与者联合行动空间中的指数爆炸的原则方法。该方法通过一个示例进行了演示,该示例显示了相对于朴素枚举的效率提升。
We present a new method for learning good strategies in zero-sum Markov games in which each side is composed of multiple agents collaborating against an opposing team of agents. Our method requires full observability and communication during learning, but the learned policies can be executed in a distributed manner. The value function is represented as a factored linear architecture and its structure determines the necessary computational resources and communication bandwidth. This approach permits a tradeoff between simple representations with little or no communication between agents and complex, computationally intensive representations with extensive coordination between agents. Thus, we provide a principled means of using approximation to combat the exponential blowup in the joint action space of the participants. The approach is demonstrated with an example that shows the efficiency gains over naive enumeration.