Contextual Bandits with Continuous Actions: Smoothing, Zooming, and Adapting

Contextual Bandits with Continuous Actions: Smoothing, Zooming, and Adapting
复制标题

具有连续动作的上下文强盗:平滑、缩放和适应

DOI:
--
复制
发表时间:
2019
期刊:
Annual Conference Computational Learning Theory
影响因子:
--
通讯作者:
Chicheng Zhang
Chicheng Zhang
中科院分区:
--
文献类型:
--
作者:
A. Krishnamurthy;J. Langford;Aleksandrs Slivkins;Chicheng Zhang

文献摘要

参考文献

被引文献

相似文献

我们研究了具有抽象策略类和连续动作空间的情境强盗学习。我们得到了两个本质上不同的后悔边界:一个在没有连续性假设的情况下与策略类的光滑版本竞争,而另一个则需要标准的Lipschitz假设。这两个界限都表现出依赖于数据的“缩放”行为,并且在没有调整的情况下,为良性问题提供了更好的保证。我们还研究了对未知光滑度参数的适应,建立了自适应代价,并推导出了不需要额外信息的最优自适应算法。
We study contextual bandit learning with an abstract policy class and continuous action space. We obtain two qualitatively different regret bounds: one competes with a smoothed version of the policy class under no continuity assumptions, while the other requires standard Lipschitz assumptions. Both bounds exhibit data-dependent "zooming" behavior and, with no tuning, yield improved guarantees for benign problems. We also study adapting to unknown smoothness parameters, establishing a price-of-adaptivity and deriving optimal adaptive algorithms that require no additional information.
DOI: --
发表时间: 2018
期刊: Proceedings of the 21st International Conference on Artificial Intelligence and Statistics (AISTATS
影响因子: --
作者:
Kallus, Nathan;Zhou, Angela
通讯作者: Zhou, Angela
自适应树强盗
DOI: 10.3150/14-bej644
发表时间: 2015
期刊: Bernoulli
影响因子: 1.5
作者:
Bull A
通讯作者: Bull A
DOI: 10.1145/3299873
发表时间: 2019-08-01
期刊: JOURNAL OF THE ACM
影响因子: 2.5
作者:
Kleinberg, Robert;Slivkins, Aleksandrs;Upfal, Eli
通讯作者: Upfal, Eli