Mean-variance and value at risk in multi-armed bandit problems
Mean-variance and value at risk in multi-armed bandit problems
复制标题
多臂老虎机问题中的均值方差和风险值
DOI:
10.1109/allerton.2015.7447162
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Qing Zhao
中科院分区:
文献类型:
--
作者:
Sattar Vakili;Qing Zhao
We study risk-averse multi-armed bandit problems under different risk measures. We consider three risk mitigation models. In the first model, the variations in the reward values obtained at different times are considered as risk and the objective is to minimize the mean-variance of the observed rewards. In the second and the third models, the quantity of interest is the total reward at the end of the time horizon, and the objective is to minimize the mean-variance and maximize the value at risk of the total reward, respectively. We develop risk-averse online learning policies and analyze their regret performance. We also provide tight lower bounds on regret under the model of mean-variance of observations.