Linear bandits with limited adaptivity and learning distributional optimal design
Linear bandits with limited adaptivity and learning distributional optimal design
复制标题
具有有限适应性和学习分布优化设计的线性老虎机
DOI:
10.1145/3406325.3451004
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Zhou, Yuan
中科院分区:
文献类型:
--
作者:
Ruan, Yufei;Yang, Jiaqi;Zhou, Yuan
Motivated by practical needs such as large-scale learning, we study the impact of adaptivity constraints to linear contextual bandits, a central problem in online learning and decision making. We consider two popular limited adaptivity models in literature: batch learning and rare policy switches. We show that, when the context vectors are adversarially chosen ind-dimensional linear contextual bandits, the learner needsO(dlogdlogT) policy switches to achieve the minimax-optimal regret, and this is optimal up topoly(logd, loglogT) factors; for stochastic context vectors, even in the more restricted batch learning model, onlyO(loglogT) batches are needed to achieve the optimal regret. Together with the known results in literature, our results present a complete picture about the adaptivity constraints in linear contextual bandits. Along the way, we propose the distributional optimal design, a natural extension of the optimal experiment design, and provide a both statistically and computationally efficient learning algorithm for the problem, which may be of independent interest.
登录
查看更多内容
影响因子:
6
作者:
Yining Wang;Adams Wei Yu;Aarti Singh
通讯作者:
Aarti Singh
DOI:
--
发表时间:
2016
期刊:
International Conference on Artificial Intelligence and Statistics
影响因子:
--
作者:
Kwang;Kevin G. Jamieson;R. Nowak;Xiaojin Zhu
通讯作者:
Xiaojin Zhu
DOI:
--
发表时间:
2018
期刊:
ACM-SIAM Symposium on Discrete Algorithms
影响因子:
--
作者:
Mohit Singh;Weijun Xie
通讯作者:
Weijun Xie
DOI:
--
发表时间:
2019
期刊:
Neural Information Processing Systems
影响因子:
--
作者:
D. Simchi;Yunzong Xu
通讯作者:
Yunzong Xu
DOI:
--
发表时间:
2015
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
作者:
Z. Guo;E. Brunskill
通讯作者:
E. Brunskill