Low-rank Tensor Bandits
Low-rank Tensor Bandits
复制标题
低阶张量强盗
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
W. Sun
中科院分区:
文献类型:
--
作者:
Botao Hao;Jie Zhou;Zheng Wen;W. Sun
In recent years, multi-dimensional online decision making has been playing a crucial role in many practical applications such as online recommendation and digital marketing. To solve it, we introduce stochastic low-rank tensor bandits, a class of bandits whose mean rewards can be represented as a low-rank tensor. We propose two learning algorithms, tensor epoch-greedy and tensor elimination, and develop finite-time regret bounds for them. We observe that tensor elimination has an optimal dependency on the time horizon, while tensor epoch-greedy has a sharper dependency on tensor dimensions. Numerical experiments further back up these theoretical findings and show that our algorithms outperform various state-of-the-art approaches that ignore the tensor low-rank structure.
影响因子:
2.5
作者:
Ming Yuan;Cun-Hui Zhang
通讯作者:
Ming Yuan;Cun-Hui Zhang