The Cost of Denied Observation in Multiagent Submodular Optimization

The Cost of Denied Observation in Multiagent Submodular Optimization
复制标题

DOI:
10.1109/cdc42340.2020.9304054
复制
发表时间:
2020-09
期刊:
2020 59th IEEE Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
D. Grimsman;Joshua H. Seaton;Jason R. Marden;Philip N. Brown
D. Grimsman;Joshua H. Seaton;Jason R. Marden;Philip N. Brown
中科院分区:
其他
文献类型:
--
作者:
D. Grimsman;Joshua H. Seaton;Jason R. Marden;Philip N. Brown

文献摘要

被引文献

相似文献

多智能体控制的一种流行形式主义应用了博弈论的工具,将多智能体决策问题转换为合作式博弈,在这种博弈中,个体智能体做出局部选择,以优化自己的局部效用函数,以响应其他智能体做出的可观察选择。当系统级目标是次模最大化时,我们知道如果每个代理人都能观察到其他所有代理人的行动选择,那么一大类结果博弈的所有纳什均衡都在最优的2倍之内;也就是说,无政府状态的代价是1/2。然而,很少有人知道,如果代理不能观察到其他相关代理的行动选择。为了研究这一点,我们扩展了标准的博弈论模型,其中一个子集的代理要么成为盲目的(无法观察别人的选择)或孤立的(盲目的,也看不见其他代理),我们证明了确切的表达式的价格无政府状态作为妥协代理的数量的函数。当k个代理商妥协时(在盲或隔离的任何组合中),我们证明了一大类效用函数的无政府状态的代价正好是1/(2 + k)。然后,我们表明,如果代理使用边际成本效用函数和至少1个妥协代理是盲目的(而不是孤立的),无政府状态的价格提高到1/(1 + k)。我们还提供了模拟结果,展示了这些观察拒绝在动态设置的影响。
A popular formalism for multiagent control applies tools from game theory, casting a multiagent decision problem as a cooperation-style game in which individual agents make local choices to optimize their own local utility functions in response to the observable choices made by other agents. When the system-level objective is submodular maximization, it is known that if every agent can observe the action choice of all other agents, then all Nash equilibria of a large class of resulting games are within a factor of 2 of optimal; that is, the price of anarchy is 1/2. However, little is known if agents cannot observe the action choices of other relevant agents. To study this, we extend the standard game-theoretic model to one in which a subset of agents either become blind (unable to observe others’ choices) or isolated (blind, and also invisible to other agents), and we prove exact expressions for the price of anarchy as a function of the number of compromised agents. When k agents are compromised (in any combination of blind or isolated), we show that the price of anarchy for a large class of utility functions is exactly 1/(2 + k). We then show that if agents use marginal-cost utility functions and at least 1 of the compromised agents is blind (rather than isolated), the price of anarchy improves to 1/(1 + k). We also provide simulation results demonstrating the effects of these observation denials in a dynamic setting.