Classification Bandits: Classification Using Expected Rewards as Imperfect Discriminators
Classification Bandits: Classification Using Expected Rewards as Imperfect Discriminators
复制标题
分类强盗:使用预期奖励作为不完美判别器进行分类
DOI:
10.1007/978-3-030-75015-2_6
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Komatsuzaki Tamiki
中科院分区:
文献类型:
--
作者:
Tabata Koji;Nakumura Atsuyoshi;Komatsuzaki Tamiki
A classification bandits problem is a new class of multi-armed bandits problems in which an agent must classify a given set of arms into positive or negative depending on whether the number of bad arms are at leastor at mostby drawing as fewer arms as possible. In our problem setting, bad arms are imperfectly characterized as the arms with above-threshold expected rewards (losses). We develop a method of reducing classification bandits to simpler one threshold classification bandits and propose an algorithm for the problem that classifies a given set of arms correctly with a specified confidence. Our numerical experiments demonstrate effectiveness of our proposed method.