A Robust Solution Concept for Human Bounded Rationality in Security Games

A Robust Solution Concept for Human Bounded Rationality in Security Games
复制标题

安全博弈中人类有限理性的鲁棒解决方案概念

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Sarit Kraus
Sarit Kraus
中科院分区:
--
文献类型:
--
作者:
J. Pita;R. John;R. Maheswaran;Milind Tambe;Sarit Kraus

文献摘要

被引文献

相似文献

博弈论的方法已被提出来解决分配有限的安全资源,以保护一组关键的目标的复杂问题。然而,许多标准假设未能解决安全部队可能面临的人类对手。为了应对这一挑战,以前的研究试图将人类决策模型整合到安全设置的博弈论算法中。目前领先的方法,实验评估的基础上,是来自一个良好的解决方案的概念,称为量子响应,被称为BRQR。一般来说,对手建模的一个关键困难是,在安全领域,有关潜在对手的信息往往是稀疏或嘈杂的,此外,游戏本身是高度复杂和规模大。因此,我们选择研究一种全新的方法来解决人类对手,避免模拟人类决策的复杂任务。我们利用和修改强大的优化技术,以创建一种新的优化类型,其中防御者的潜在偏差的攻击者的损失是有界的距离偏离预期值最大化的策略。为了证明我们的方法的优势,我们引入了一个系统的方法来生成有意义的奖励结构,并比较我们的方法与BRQR在最全面的调查,迄今为止,涉及104个安全设置,以前的工作只测试了10个安全设置。我们的实验分析表明,我们的方法执行以及或优于BRQR在超过90%的安全设置测试,我们表现出显着的运行时的好处。这些结果有利于在这些复杂领域中利用基于鲁棒优化的方法来避免对手建模的困难。
Game-theoretic approaches have been proposed for addressing the complex problem of assigning limited security resources to protect a critical set of targets. However, many of the standard assumptions fail to address human adversaries who security forces will likely face. To address this challenge, previous research has attempted to integrate models of human decision-making into the game-theoretic algorithms for security settings. The current leading approach, based on experimental evaluation, is derived from a well-founded solution concept known as quantal response and is known as BRQR. One critical difficulty with opponent modeling in general is that, in security domains, information about potential adversaries is often sparse or noisy and furthermore, the games themselves are highly complex and large in scale. Thus, we chose to examine a completely new approach to addressing human adversaries that avoids the complex task of modeling human decision-making. We leverage and modify robust optimization techniques to create a new type of optimization where the defender’s loss for a potential deviation by the attacker is bounded by the distance of that deviation from the expected-value-maximizing strategy. To demonstrate the advantages of our approach, we introduce a systematic way to generate meaningful reward structures and compare our approach with BRQR in the most comprehensive investigation to date involving 104 security settings where previous work has tested only up to 10 security settings. Our experimental analysis reveals our approach performing as well as or outperforming BRQR in over 90% of the security settings tested and we demonstrate significant runtime benefits. These results are in favor of utilizing an approach based on robust optimization in these complex domains to avoid the difficulties of opponent modeling.