Classification Bandits: Classification Using Expected Rewards as Imperfect Discriminators

Classification Bandits: Classification Using Expected Rewards as Imperfect Discriminators
复制标题

分类强盗:使用预期奖励作为不完美判别器进行分类

DOI:
10.1007/978-3-030-75015-2_6
复制
发表时间:
2021
期刊:
Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), Workshop on Machine Learning for MEasurement INformatics
影响因子:
--
通讯作者:
Komatsuzaki Tamiki
Komatsuzaki Tamiki
中科院分区:
--
文献类型:
--
作者:
Tabata Koji;Nakumura Atsuyoshi;Komatsuzaki Tamiki

文献摘要

相似文献

分类强盗问题是一类新的多武器强盗问题,其中智能体必须根据不良武器的数量是否最少或最多,通过尽可能少地画出武器,将给定的一组武器分类为正或负。在我们的问题设置中,坏武器被不完美地描述为具有高于阈值的预期奖励(损失)的武器。我们开发了一种方法,减少分类强盗简单的一个阈值分类强盗,并提出了一个算法的问题,正确地分类一组给定的武器与指定的信心。我们的数值实验证明了我们所提出的方法的有效性。
A classification bandits problem is a new class of multi-armed bandits problems in which an agent must classify a given set of arms into positive or negative depending on whether the number of bad arms are at leastor at mostby drawing as fewer arms as possible. In our problem setting, bad arms are imperfectly characterized as the arms with above-threshold expected rewards (losses). We develop a method of reducing classification bandits to simpler one threshold classification bandits and propose an algorithm for the problem that classifies a given set of arms correctly with a specified confidence. Our numerical experiments demonstrate effectiveness of our proposed method.