Optimal Adaptive Policies for Sequential Allocation Problems
Optimal Adaptive Policies for Sequential Allocation Problems
复制标题
顺序分配问题的最优自适应策略
DOI:
--
复制
发表时间:
1996
期刊:
影响因子:
--
通讯作者:
M. Katehakis
中科院分区:
文献类型:
--
作者:
A. Burnetas;M. Katehakis
Consider the problem of sequential sampling frommstatistical populations to maximize the expected sum of outcomes in the long run. Under suitable assumptions on the unknown parametersformula, it is shown that there exists a classCRof adaptive policies with the following properties: (i) The expectednhorizon rewardformulaunder any policy Â?0inCRis equal toformula, asnÂ?∞, whereformulais the largest population mean andformulais a constant. (ii) Policies inCRare asymptotically optimal within a larger classCUFof “uniformly fast convergent” policies in the sense thatformula, for any Â?Â?CUFand anyformulasuch thatformula. Policies inCRare specified via easily computable indices, defined as unique solutions to dual problems that arise naturally from the functional form offormula. In addition, the assumptions are verified for populations specified by nonparametric discrete univariate distributions with finite support. In the case of normal populations with unknown means and variances, we leave as an open problem the verification of one assumption.