Generalized Linear Bandits with Safety Constraints

Generalized Linear Bandits with Safety Constraints
复制标题

具有安全约束的广义线性老虎机

DOI:
10.1109/icassp40776.2020.9054063
复制
发表时间:
2020
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Christos Thrampoulidis
Christos Thrampoulidis
中科院分区:
--
文献类型:
--
作者:
Sanae Amani;M. Alizadeh;Christos Thrampoulidis

文献摘要

被引文献

相似文献

经典的多臂强盗是一类序贯决策问题,其中选择行动的代价是从未知的潜在分布中独立抽样的。Bandit算法在安全关键系统中有许多应用,其中尽管问题参数是不确定的,但在算法的运行过程中必须遵守几个约束。本文建立了一个广义线性随机多臂强盗问题,该问题具有依赖于未知参数向量的广义线性安全约束。在这种情况下,我们提出了一个安全的UCB-GLM算法,我们给出了一般的和问题相关的后悔边界。
The classical multi-armed bandit is a class of sequential decision making problems where selecting actions incurs costs that are sampled independently from an unknown underlying distribution. Bandit algorithms have many applications in safety critical systems, where several constraints must be respected during the run of the algorithm in spite of uncertainty about problem parameters. This paper formulates a generalized linear stochastic multi-armed bandit problem with generalized linear safety constraints that depend on an unknown parameter vector. In this setting, we propose a Safe UCB-GLM algorithm for which we provide general and problem-dependent regret bounds.