Negatively Correlated Bandits

Negatively Correlated Bandits
复制标题

DOI:
10.1093/restud/rdq025
复制
发表时间:
2008-10
期刊:
Microeconomics: Search; Learning; Information Costs & Specific Knowledge; Expectation & Speculation eJournal
影响因子:
--
通讯作者:
N. Klein;Sven Rady
N. Klein;Sven Rady
中科院分区:
其他
文献类型:
--
作者:
N. Klein;Sven Rady

文献摘要

被引文献

相似文献

我们分析了双臂强盗的两人战略实验游戏。每个玩家必须在连续的时间内决定是否使用具有已知收益的安全策略或最初未知提供收益的可能性的风险策略。玩家之间风险武器的质量完全负相关。与两个风险臂都属于同一类型的情况形成鲜明对比的是,我们发现,如果赌注超过某个阈值,则在任何马尔可夫完美均衡中学习都将完成,并且所有均衡都处于截止策略中。对于低风险,均衡是唯一的、对称的,并且与规划者的解决方案一致。对于高风险来说,平衡是独特的、对称的,并且等同于短视行为。对于中间赌注来说,存在一个连续的均衡。
We analyze a two-player game of strategic experimentation with two-armed bandits. Each player has to decide in continuous time whether to use a safe arm with a known payoff or a risky arm whose likelihood of delivering payoffs is initially unknown. The quality of the risky arms is perfectly negatively correlated between players. In marked contrast to the case where both risky arms are of the same type, we find that learning will be complete in any Markov perfect equilibrium if the stakes exceed a certain threshold, and that all equilibria are in cutoff strategies. For low stakes, the equilibrium is unique, symmetric, and coincides with the planner's solution. For high stakes, the equilibrium is unique, symmetric, and tantamount to myopic behavior. For intermediate stakes, there is a continuum of equilibria.