Optimal Off-line Experimentation for Games

Optimal Off-line Experimentation for Games
复制标题

DOI:
10.1287/deca.2020.0412
复制
发表时间:
2020-05
期刊:
Decis. Anal.
影响因子:
--
通讯作者:
T. Allen;Olivia K. Hernandez;A. Alomair
T. Allen;Olivia K. Hernandez;A. Alomair
中科院分区:
其他
文献类型:
--
作者:
T. Allen;Olivia K. Hernandez;A. Alomair

文献摘要

相似文献

许多商业情况可以被称为“游戏”,因为结果取决于具有不同目标的多个决策者。然而,在许多情况下,玩家选项的所有组合的收益是不可用的,但离线实验的能力是可用的。例如,战争游戏演习,测试营销,网络范围的活动,以及许多类型的模拟都可以被视为离线游戏相关的实验。我们解决的决策问题,规划和分析离线实验的游戏与初始程序,以尽量减少错误的回报估计。然后,我们提供了一个顺序的算法,减少选择的选项组合是无关的候选人纳什,相关的,累积前景理论或其他均衡的评估。我们还提供了一个有效的公式来估计的机会,给定的纳什均衡存在,提供有关一般均衡的收敛保证,并提供了一个停止标准称为估计的期望值完美的离线信息(EEVPOI)。EEVPOI是基于进一步离线实验的预期效用的有界增益。基于网络安全夺旗游戏,提供了使用仿真模型来说明所提出的所有方法的例子。该例子表明,所提出的方法,使大量减少的测试运行(一半)的数量相比,一个完整的阶乘和停止标准的计算时间。
Many business situations can be called “games” because outcomes depend on multiple decision makers with differing objectives. Yet, in many cases, the payoffs for all combinations of player options are not available, but the ability to experiment off-line is available. For example, war-gaming exercises, test marketing, cyber-range activities, and many types of simulations can all be viewed as off-line gaming-related experimentation. We address the decision problem of planning and analyzing off-line experimentation for games with an initial procedure seeking to minimize the errors in payoff estimates. Then, we provide a sequential algorithm with reduced selections from option combinations that are irrelevant to evaluating candidate Nash, correlated, cumulative prospect theory or other equilibria. We also provide an efficient formula to estimate the chance that given Nash equilibria exists, provide convergence guarantees relating to general equilibria, and provide a stopping criterion called the estimated expected value of perfect off-line information (EEVPOI). The EEVPOI is based on bounded gains in expected utility from further off-line experimentation. An example of using a simulation model to illustrate all the proposed methods is provided based on a cyber security capture-the-flag game. The example demonstrates that the proposed methods enable substantial reductions in both the number of test runs (half) compared with a full factorial and the computational time for the stopping criterion.