Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits

Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits
复制标题

DOI:
10.1609/aaai.v35i9.16961
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yu-Heng Hung;Ping-Chun Hsieh;Xi Liu;P. Kumar
Yu-Heng Hung;Ping-Chun Hsieh;Xi Liu;P. Kumar
中科院分区:
其他
文献类型:
--
作者:
Yu-Heng Hung;Ping-Chun Hsieh;Xi Liu;P. Kumar

文献摘要

被引文献

相似文献

修改奖励偏置最大似然方法最初提出的自适应控制文献中,我们提出了新的学习算法来处理探索利用权衡线性土匪问题以及广义线性土匪问题。我们开发了新的指数政策,我们证明实现了订单最优,并表明他们实现了经验的表现与国家的最先进的基准方法在广泛的实验竞争。新的政策实现了这一点,每次拉低计算时间的线性土匪,从而导致有利的遗憾,以及计算效率。
Modifying the reward-biased maximum likelihood method originally proposed in the adaptive control literature, we propose novel learning algorithms to handle the explore-exploit trade-off in linear bandits problems as well as generalized linear bandits problems. We develop novel index policies that we prove achieve order-optimality, and show that they achieve empirical performance competitive with the state-of-the-art benchmark methods in extensive experiments. The new policies achieve this with low computation time per pull for linear bandits, and thereby resulting in both favorable regret as well as computational efficiency.