Risk-Aware Multi-Armed Bandits With Refined Upper Confidence Bounds

Risk-Aware Multi-Armed Bandits With Refined Upper Confidence Bounds
复制标题

具有精细置信上限的具有风险意识的多臂强盗

DOI:
--
复制
发表时间:
2021
影响因子:
3.9
通讯作者:
M. van der Schaar
M. van der Schaar
中科院分区:
工程技术2区
文献类型:
--
作者:
Xingchi Liu;Mahsa Derakhshani;S. Lambotharan;M. van der Schaar

文献摘要

被引文献

相似文献

经典的多臂强盗(MAB)框架研究了决策问题的探索-开发困境,总是将具有最高期望回报的手臂视为最优选择。然而,在某些应用中,如果方差很高,则具有高期望奖励的手臂可能会有风险。因此,应该考虑报酬的变化,使手臂选择过程具有风险意识。在这封信中,均值方差度量的研究,以衡量收到的奖励的不确定性。首先研究了收益服从高斯分布时的风险感知MAB问题,利用方差的集中不等式设计了一个高斯风险感知置信上界算法.此外,我们将该算法扩展到一个新的渐近风险意识的置信上界算法,通过开发一个上界的方差的样本方差的渐近分布的基础上。理论分析证明,这两种算法都达到了$mathcal {O}(log(T))$遗憾.最后,数值结果表明,我们的算法优于几个风险意识的MAB算法。
The classical multi-armed bandit (MAB) framework studies the exploration-exploitation dilemma of the decision-making problem and always treats the arm with the highest expected reward as the optimal choice. However, in some applications, an arm with a high expected reward can be risky to play if the variance is high. Hence, the variation of the reward should be considered to make the arm-selection process risk-aware. In this letter, the mean-variance metric is investigated to measure the uncertainty of the received rewards. We first study a risk-aware MAB problem when the reward follows a Gaussian distribution, and a concentration inequality on the variance is developed to design a Gaussian risk aware-upper confidence bound algorithm. Furthermore, we extend this algorithm to a novel asymptotic risk aware-upper confidence bound algorithm by developing an upper confidence bound of the variance based on the asymptotic distribution of the sample variance. Theoretical analysis proves that both proposed algorithms achieve the $mathcal {O}(log (T))$ regret. Finally, numerical results demonstrate that our algorithms outperform several risk-aware MAB algorithms.