EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits

EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yikun Ban;Yuchen Yan;A. Banerjee;Jingrui He
Yikun Ban;Yuchen Yan;A. Banerjee;Jingrui He
中科院分区:
其他
文献类型:
--
作者:
Yikun Ban;Yuchen Yan;A. Banerjee;Jingrui He

文献摘要

相似文献

在本文中,我们在情境老虎机中提出了一种新颖的神经探索策略——EE - Net,它不同于基于上置信界(UCB)和汤普森采样(TS)的标准方法。情境多臂老虎机已经被研究了数十年,有着各种各样的应用。为了解决老虎机中的利用 - 探索权衡问题,主要有三种技术:epsilon - 贪婪算法、汤普森采样(TS)和上置信界(UCB)。在近期的文献中,线性情境老虎机采用岭回归来估计奖励函数,并将其与TS或UCB策略相结合用于探索。然而,这类工作明确假设奖励是基于臂向量的线性函数,这在现实世界的数据集中可能并不成立。为了克服这一挑战,一系列神经老虎机算法被提出,其中使用神经网络来学习潜在的奖励函数,并对TS或UCB进行调整以用于探索。我们提出“EE - Net”这种新颖的基于神经的探索策略,而不是像先前的方法那样计算基于大偏差的统计界用于探索。除了使用一个神经网络(利用网络)来学习奖励函数外,EE - Net还使用另一个神经网络(探索网络)来自适应地学习与当前估计奖励相比的潜在收益以用于探索。然后,构建一个决策器来组合利用网络和探索网络的输出。我们证明EE - Net能够实现$\mathcal{O}(\sqrt{T\log T})$的遗憾值,并表明EE - Net在现实世界的数据集中优于现有的线性和神经情境老虎机基准算法。
In this paper, we propose a novel neural exploration strategy in contextual bandits, EE-Net, distinct from the standard UCB-based and TS-based approaches. Contextual multi-armed bandits have been studied for decades with various applications. To solve the exploitation-exploration tradeoff in bandits, there are three main techniques: epsilon-greedy, Thompson Sampling (TS), and Upper Confidence Bound (UCB). In recent literature, linear contextual bandits have adopted ridge regression to estimate the reward function and combine it with TS or UCB strategies for exploration. However, this line of works explicitly assumes the reward is based on a linear function of arm vectors, which may not be true in real-world datasets. To overcome this challenge, a series of neural bandit algorithms have been proposed, where a neural network is used to learn the underlying reward function and TS or UCB are adapted for exploration. Instead of calculating a large-deviation based statistical bound for exploration like previous methods, we propose"EE-Net", a novel neural-based exploration strategy. In addition to using a neural network (Exploitation network) to learn the reward function, EE-Net uses another neural network (Exploration network) to adaptively learn potential gains compared to the currently estimated reward for exploration. Then, a decision-maker is constructed to combine the outputs from the Exploitation and Exploration networks. We prove that EE-Net can achieve $\mathcal{O}(\sqrt{T\log T})$ regret and show that EE-Net outperforms existing linear and neural contextual bandit baselines on real-world datasets.