Bandit problems with Levy payoff processes

Bandit problems with Levy payoff processes
复制标题

征费支付流程的强盗问题

DOI:
--
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
Eilon Solan
Eilon Solan
中科院分区:
--
文献类型:
--
作者:
A. Cohen;Eilon Solan

文献摘要

被引文献

相似文献

我们研究连续时间的双臂 Levy 老虎机,其中一个安全臂产生恒定的收益 s,另一个风险臂可以是高类型或低类型;这两种类型都会产生由征税过程产生的随机收益。当臂为高时,Levy 过程的期望大于 s,如果臂为低,则低于 s。 决策者 (DM) 必须在任何给定时间 t 选择在时间间隔 [t,t+dt) 内分配给每个臂的资源比例。我们证明,在 Levy 过程的适当条件下,存在唯一的最优策略,即截止策略,并且我们提供了截止和最优收益的显式公式,作为问题数据的函数。我们还检查了 DM 对风险臂类型的先验不正确的情况,并计算了 DM 采取与不正确的先验相对应的最优策略所获得的预期收益。 此外,我们研究了结果的两种应用:(a)我们展示了如何在双臂 Levy bandit 问题中对信息进行定价,(b)我们研究谁在双臂 Levy bandit 问题中表现更好:乐观主义者分配给 High 的概率高于真实概率,或者悲观主义者分配给 High 的概率低于真实概率。
We study two-armed Levy bandits in continuous-time, which have one safe arm that yields a constant payoff s, and one risky arm that can be either of type High or Low; both types yield stochastic payoffs generated by a Levy process. The expectation of the Levy process when the arm is High is greater than s, and lower than s if the arm is Low. The decision maker (DM) has to choose, at any given time t, the fraction of resource to be allocated to each arm over the time interval [t,t+dt). We show that under proper conditions on the Levy processes, there is a unique optimal strategy, which is a cut-off strategy, and we provide an explicit formula for the cut-off and the optimal payoff, as a function of the data of the problem. We also examine the case where the DM has incorrect prior over the type of the risky arm, and we calculate the expected payoff gained by a DM who plays the optimal strategy that corresponds to the incorrect prior. In addition, we study two applications of the results: (a) we show how to price information in two-armed Levy bandit problem, and (b) we investigate who fares better in two-armed bandit problems: an optimist who assigns to High a probability higher than the true probability, or a pessimist who assigns to High a probability lower than the true probability.