Sample Efficient Multiagent Learning in the Presence of Markovian Agents

Sample Efficient Multiagent Learning in the Presence of Markovian Agents
复制标题

在马尔可夫智能体存在的情况下进行高效多智能体学习的示例

DOI:
--
复制
发表时间:
2013
期刊:
Studies in Computational Intelligence
影响因子:
--
通讯作者:
D. Chakraborty
D. Chakraborty
中科院分区:
--
文献类型:
--
作者:
D. Chakraborty

文献摘要

被引文献

相似文献

多智能体学习(或MAL)的问题是关于研究在其他智能体同时适应的情况下,智能体如何学习和适应。这个问题经常在重复矩阵博弈提供的程式化设置中进行研究。这本书的目标是为这样的设置开发MAL算法,以实现以前没有实现的一组新目标。本书有三个主要贡献。第一个主要贡献提出了一种新的MAL算法,称为收敛与模型学习和安全(或CMLeS),这是第一个实现以下三个目标:(1)收敛到以下纳什均衡联合策略的自我发挥:(2)实现接近最佳的反应时,与一组记忆有限的代理,其内存大小的上限由一个已知的值;以及(3)确保在与任何其他代理集合交互时非常接近其安全值的个体回报。第二个主要贡献提出了另一种新的MAL算法,该算法模拟了一类更为复杂的代理行为,称为马尔可夫代理,它包含了一类内存有界的代理。它被称为针对马尔可夫代理的联合优化(或Joma),它实现了以下两个目标:(1)当与马尔可夫代理交互时,实现非常接近社会福利最大化联合回报的联合回报;(2)当与任何其他代理集交互时,确保个人回报非常接近其安全值。最后,第三个主要贡献展示了如何扩展Joma的一个关键子程序,以解决与强化学习有关的更广泛的一类问题,称为“因子化状态MDP中的结构学习”。本书中提出的所有算法都有严格的理论分析作为支持,包括对样本复杂性的分析,以及代表性的经验测试。德克萨斯大学奥斯汀分校Doran Chakraborty 2013主管:Peter Stone
The problem of multiagent learning (or MAL) is concerned with the study of how agents can learn and adapt in the presence of other agents that are simultaneously adapting. The problem is often studied in the stylized settings provided by repeated matrix games. The goal of this book is to develop MAL algorithms for such a setting that achieve a new set of objectives which have not been previously achieved. The book makes three main contributions. The first main contribution proposes a novel MAL algorithm, called Convergence with Model Learning and Safety (or CMLeS), that is the first to achieve the following three objectives: (1) converges to following a Nash equilibrium joint-policy in self-play; (2) achieves close to the best response when interacting with a set of memory-bounded agents whose memory size is upper bounded by a known value; and (3) ensures an individual return that is very close to its security value when interacting with any other set of agents. The second main contribution proposes another novel MAL algorithm that models a significantly more complex class of agent behavior called Markovian agents, that subsumes the class of memory-bounded agents. Called Joint Optimization against Markovian Agents (or Joma), it achieves the following two objectives: (1) achieves a joint-return very close to the social welfare maximizing joint-return when interacting with Markovian agents; (2) ensures an individual return that is very close to its security value when interacting with any other set of agents. Finally, the third main contribution shows how a key subroutine of Joma can be extended to solve a broader class of problems pertaining to Reinforcement Learning, called “Structure Learning in factored state MDPs”. All of the algorithms presented in this book are well backed with rigorous theoretical analysis, including an analysis on sample complexity wherever applicable, as well as representative empirical tests. University of Texas, Austin Doran Chakraborty 2013 Supervisor: Peter Stone