An experimental analysis of the bandit problem

An experimental analysis of the bandit problem
复制标题

DOI:
10.1007/s001990050146
复制
发表时间:
1997-06-01
期刊:
影响因子:
1.3
通讯作者:
Porter, D
Porter, D
中科院分区:
经济学3区
文献类型:
--
作者:
Banks, J;Olson, M;Porter, D

文献摘要

被引文献

相似文献

我们调查,在一个实验环境中,行为的单个决策者在离散的时间间隔在一个“无限”的地平线可能会选择一个行动从一组可能的行动,这组是恒定的,随着时间的推移,即一个强盗问题。两个强盗环境进行了研究,其中一个预测的行为应该总是近视(双臂强盗),另一个预测的行为应该永远不会近视(单臂强盗)。我们还调查了比较静态的预测作为潜在的参数的强盗环境的变化。综合结果表明,在两个强盗环境中的行为是定量不同的,在理论预测的方向。
We investigate, in an experimental setting, the behavior of single decision makers who at discrete time intervals over an ''infinite'' horizon may choose one action from a set of possible actions where this set is constant over time, i.e. a bandit problem. Two bandit environments an examined, one in which the predicted behavior should always be myopic (the two-armed bandit) and the other in which the predicted behavior should never be myopic (the one-armed bandit). We also investigate the comparative static predictions as the underlying parameters of the bandit environments are changed. The aggregate results show that the behavior in the two bandit environments are quantitatively different and in the direction of the theoretical predictions.