The Knowledge Gradient Algorithm for a General Class of Online Learning Problems

The Knowledge Gradient Algorithm for a General Class of Online Learning Problems
复制标题

DOI:
10.1287/opre.1110.0999
复制
发表时间:
2012
期刊:
Oper. Res.
影响因子:
--
通讯作者:
I. Ryzhov;Warren B. Powell;Peter I. Frazier
I. Ryzhov;Warren B. Powell;Peter I. Frazier
中科院分区:
其他
文献类型:
--
作者:
I. Ryzhov;Warren B. Powell;Peter I. Frazier

文献摘要

被引文献

相似文献

我们推导出一个一个时期的前瞻政策有限和无限地平线在线最优学习问题的高斯奖励。我们的方法是能够处理的情况下,我们的先验信念的奖励是相关的,这是不处理传统的多臂强盗方法。实验表明,我们的KG政策执行竞争对手的最佳策略在经典的强盗问题的最佳逼近,它优于许多学习政策的相关情况下。
We derive a one-period look-ahead policy for finite-and infinite-horizon online optimal learning problems with Gaussian rewards. Our approach is able to handle the case where our prior beliefs about the rewards are correlated, which is not handled by traditional multiarmed bandit methods. Experiments show that our KG policy performs competitively against the best-known approximation to the optimal policy in the classic bandit problem, and it outperforms many learning policies in the correlated case.