Design of Experiments for Stochastic Contextual Linear Bandits

Design of Experiments for Stochastic Contextual Linear Bandits
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
A. Zanette;Kefan Dong;Jonathan Lee;E. Brunskill
A. Zanette;Kefan Dong;Jonathan Lee;E. Brunskill
中科院分区:
其他
文献类型:
--
作者:
A. Zanette;Kefan Dong;Jonathan Lee;E. Brunskill

文献摘要

相似文献

在随机线性上下文的强盗设置存在几个极大极小的探索过程与政策是反应的数据被收购。在实践中,部署这些算法可能会带来巨大的工程开销,特别是当数据集以分布式方式收集时,或者当需要人工参与来实现不同的策略时。在这种情况下,使用单个非反应策略进行探索是有益的。假设一些批处理上下文是可用的,我们设计了一个单一的随机策略来收集一个好的数据集,从中可以提取一个接近最优的策略。我们提出了一个理论分析,以及在合成和真实世界的数据集上的数值实验。
In the stochastic linear contextual bandit setting there exist several minimax procedures for exploration with policies that are reactive to the data being acquired. In practice, there can be a significant engineering overhead to deploy these algorithms, especially when the dataset is collected in a distributed fashion or when a human in the loop is needed to implement a different policy. Exploring with a single non-reactive policy is beneficial in such cases. Assuming some batch contexts are available, we design a single stochastic policy to collect a good dataset from which a near-optimal policy can be extracted. We present a theoretical analysis as well as numerical experiments on both synthetic and real-world datasets.