Risk-Aware Multi-Armed Bandits With Refined Upper Confidence Bounds
Risk-Aware Multi-Armed Bandits With Refined Upper Confidence Bounds
复制标题
具有精细置信上限的具有风险意识的多臂强盗
DOI:
--
复制
发表时间:
2021
影响因子:
3.9
通讯作者:
M. van der Schaar
中科院分区:
文献类型:
--
作者:
Xingchi Liu;Mahsa Derakhshani;S. Lambotharan;M. van der Schaar
The classical multi-armed bandit (MAB) framework studies the exploration-exploitation dilemma of the decision-making problem and always treats the arm with the highest expected reward as the optimal choice. However, in some applications, an arm with a high expected reward can be risky to play if the variance is high. Hence, the variation of the reward should be considered to make the arm-selection process risk-aware. In this letter, the mean-variance metric is investigated to measure the uncertainty of the received rewards. We first study a risk-aware MAB problem when the reward follows a Gaussian distribution, and a concentration inequality on the variance is developed to design a Gaussian risk aware-upper confidence bound algorithm. Furthermore, we extend this algorithm to a novel asymptotic risk aware-upper confidence bound algorithm by developing an upper confidence bound of the variance based on the asymptotic distribution of the sample variance. Theoretical analysis proves that both proposed algorithms achieve the $mathcal {O}(log (T))$ regret. Finally, numerical results demonstrate that our algorithms outperform several risk-aware MAB algorithms.