Neural Contextual Bandits with UCB-based Exploration

Neural Contextual Bandits with UCB-based Exploration
复制标题

DOI:
--
复制
发表时间:
2019-11
期刊:
--
影响因子:
--
通讯作者:
Dongruo Zhou;Lihong Li;Quanquan Gu
Dongruo Zhou;Lihong Li;Quanquan Gu
中科院分区:
其他
文献类型:
--
作者:
Dongruo Zhou;Lihong Li;Quanquan Gu

文献摘要

被引文献

相似文献

我们研究随机上下文强盗问题,其中奖励是从一个未知的函数与加性噪声。除了有界性之外,没有对奖励函数做任何假设。我们提出了一种新算法NeuralUCB,它利用深度神经网络的表示能力,并使用基于神经网络的随机特征映射来构建有效探索的奖励置信上限(UCB)。我们证明,在标准假设下,NeuralUCB实现$\tilde O(\sqrt{T})$遗憾,其中$T$是轮数。据我们所知,它是第一个基于神经网络的上下文强盗算法,具有接近最优的遗憾保证。我们还表明,该算法是经验上有竞争力的代表性的基线在一些基准。
We study the stochastic contextual bandit problem, where the reward is generated from an unknown function with additive noise. No assumption is made about the reward function other than boundedness. We propose a new algorithm, NeuralUCB, which leverages the representation power of deep neural networks and uses a neural network-based random feature mapping to construct an upper confidence bound (UCB) of reward for efficient exploration. We prove that, under standard assumptions, NeuralUCB achieves $\tilde O(\sqrt{T})$ regret, where $T$ is the number of rounds. To the best of our knowledge, it is the first neural network-based contextual bandit algorithm with a near-optimal regret guarantee. We also show the algorithm is empirically competitive against representative baselines in a number of benchmarks.