Optimal Adaptive Policies for Sequential Allocation Problems

Optimal Adaptive Policies for Sequential Allocation Problems
复制标题

顺序分配问题的最优自适应策略

DOI:
--
复制
发表时间:
1996
期刊:
影响因子:
--
通讯作者:
M. Katehakis
M. Katehakis
中科院分区:
--
文献类型:
--
作者:
A. Burnetas;M. Katehakis

文献摘要

被引文献

相似文献

考虑从统计总体中进行序贯抽样以最大化长期预期结果之和的问题。在对未知参数公式的适当假设下,证明了存在一类具有以下性质的自适应策略:(i)任意策略下的期望非地平线报酬公式?0inCR is equal tofrule,as n?∞,其中公式为最大总体均值,公式为常数。(ii)CR中的政策是渐近最优的一个更大的类CUF的“一致快速收敛”的政策在这个意义上thatformula,对于任何??CUF和任何这样的公式。CR中的政策是通过容易计算的指数来指定的,定义为从公式的函数形式自然产生的对偶问题的唯一解决方案。此外,验证了假设的非参数离散单变量分布与有限的支持指定的人口。在具有未知均值和方差的正态总体的情况下,我们将一个假设的验证作为一个开放问题。
Consider the problem of sequential sampling frommstatistical populations to maximize the expected sum of outcomes in the long run. Under suitable assumptions on the unknown parametersformula, it is shown that there exists a classCRof adaptive policies with the following properties: (i) The expectednhorizon rewardformulaunder any policy Â?0inCRis equal toformula, asnÂ?∞, whereformulais the largest population mean andformulais a constant. (ii) Policies inCRare asymptotically optimal within a larger classCUFof “uniformly fast convergent” policies in the sense thatformula, for any Â?Â?CUFand anyformulasuch thatformula. Policies inCRare specified via easily computable indices, defined as unique solutions to dual problems that arise naturally from the functional form offormula. In addition, the assumptions are verified for populations specified by nonparametric discrete univariate distributions with finite support. In the case of normal populations with unknown means and variances, we leave as an open problem the verification of one assumption.