Multi-player multi-armed bandits: Decentralized learning with IID rewards

Multi-player multi-armed bandits: Decentralized learning with IID rewards
复制标题

多人多臂强盗:具有 IID 奖励的去中心化学习

DOI:
--
复制
发表时间:
2012
期刊:
Allerton Conference on Communication, Control, and Computing
影响因子:
--
通讯作者:
R. Jain
R. Jain
中科院分区:
--
文献类型:
--
作者:
D. Kalathil;Naumaan Nayyar;R. Jain

文献摘要

被引文献

相似文献

研究了分散多臂强盗问题,每个局中人有不同的武器。每个玩家可以在每个时刻选择一只手臂,并可以从具有未知均值的未知分布中获得随机奖励。武器给不同的玩家不同的奖励。如果超过一个玩家选择同一只手臂,每个人都得到零奖励。不存在用于用户之间的通信或协调的专用控制信道。我们提出了一个在线学习算法称为dUCB4,实现了近O(log2 T)。动机来自认知无线电网络中的多个次级用户的机会性频谱接入,其中他们必须在对不同用户看起来不同的各种无线信道中进行挑选。
We consider the decentralized multi-armed bandit problem with distinct arms for each players. Each player can pick one arm at each time instant and can get a random reward from an unknown distribution with an unknown mean. The arms give different rewards to different players. If more than one player select the same arm, everyone gets a zero reward. There is no dedicated control channel for communication or coordination among the user. We propose an online learning algorithm called dUCB4 which achieves a near-O(log2 T). The motivation comes from opportunistic spectrum access by multiple secondary users in cognitive radio networks wherein they must pick among various wireless channels that look different to different users.