Multi-Armed Bandits: Theory and Applications to Online Learning in Networks

Multi-Armed Bandits: Theory and Applications to Online Learning in Networks
复制标题

多臂强盗:网络在线学习的理论与应用

DOI:
10.2200/s00941ed2v01y201907cnt022
复制
发表时间:
2019
期刊:
Synthesis Lectures on Communication Networks
影响因子:
--
通讯作者:
Qing Zhao
Qing Zhao
中科院分区:
--
文献类型:
--
作者:
Qing Zhao

文献摘要

参考文献

被引文献

相似文献

摘要 多臂老虎机问题涉及未知环境中的最优顺序决策和学习。自从 1933 年 Thompson 为应用程序提出第一个强盗问题以来......
Abstract Multi-armed bandit problems pertain to optimal sequential decision making and learning in unknown environments. Since the first bandit problem posed by Thompson in 1933 for the application...
自适应树强盗
DOI: 10.3150/14-bej644
发表时间: 2015
期刊: Bernoulli
影响因子: 1.5
作者:
Bull A
通讯作者: Bull A
DOI: 10.1145/3299873
发表时间: 2019-08-01
期刊: JOURNAL OF THE ACM
影响因子: 2.5
作者:
Kleinberg, Robert;Slivkins, Aleksandrs;Upfal, Eli
通讯作者: Upfal, Eli